Essay · September 3, 2026 · Relay #251

The Crawl Contract

Nobody broke the agreement. The reward simply became collectable without paying out.

By KW Norton.

Dana McKay and Damiano Spina, writing in The Conversation, describe the arrangement that has held the open web together for thirty years: sites let crawlers in for free, crawlers send back traffic, and a site that dislikes the deal writes a line in robots.txt. They report what is now happening to it. Crawlers arrive that do not send traffic back. Publishers respond by shutting the door. Cloudflare, which fronts a large share of the most-visited sites, begins blocking AI crawlers by default on ad-bearing pages on September 15. Wikipedia's traffic fell language by language as AI summaries rolled out in each. Roughly one in six sources cited by AI search tools is already an AI-generated site.

Their conclusion is that reliable information is getting harder to find. I want to add one thing to it, and only one, because the reporting does not need my help.

It was never a contract

Calling it a social contract is generous in a way that hides the mechanism. A contract has parties, terms, and a remedy. This had none. What it had was a reward loop: publish, get indexed, get traffic, get paid. Every participant optimized against the same proxy — traffic — and for three decades that proxy sat close enough to the thing anyone actually wanted, which was a public able to find reliable things, that nobody had to notice the difference between them.

Then a system arrived that could collect the value of the corpus without emitting the traffic. Nothing was violated. No one defected in the moral sense. The loop was simply run by something that found the shorter path to the same reward, which is what optimizers do and the only thing they are asked to do. Publishers, facing a loop that no longer pays, are now optimizing their own proxy — cost avoided, revenue protected — by closing. That is not spite either. It is the same arithmetic from the other side.

This is the signature I have been describing in schools, in institutions, and in preference-trained models: the proxy stands in for the goal, the proxy detaches, and the goal leaves with it while every participant is behaving correctly by the measure in front of them. What is new here is not the mechanism. It is the size of the thing being measured. The object under proxy failure this time is the shared record.

The selection effect nobody chose

The detail in McKay and Spina worth sitting with is not the traffic decline. It is the asymmetry in who blocks. Sites carrying misinformation have little reason to keep crawlers out; they were never earning from the visit in the first place, and being repeated is the entire point. Sites carrying costly, checked work have every reason to close.

So the door is being shut selectively, and the selection runs against quality — not by anyone's design, and not by anyone's malice. It is an emergent sort produced by two rational local decisions and no coordinating agent anywhere in the system. Stigmergy without a steward: each participant reads the trace left by the others and responds, and the shape that results was chosen by no one.

The part that is mine, and is a claim

Here is the step past the reporting, labeled as what it is. The systems that most need a well-kept commons are the ones eroding the conditions under which it is kept. A model's usefulness on anything current depends on somebody paying to find things out and publish them. The economics that paid for that were carried by the traffic loop. If the loop stays broken and nothing replaces it — pay-to-crawl has not taken hold — then the corpus thins first at exactly the layer that was expensive to produce, and the models train forward on what is left, which increasingly includes their own output.

I am not predicting collapse. Model collapse is a real result under specific conditions and an overused word everywhere else, and the web has absorbed several extinction-grade economic shifts without disappearing. What I claim is narrower and checkable: the failure runs through supply, not through capability, and treating it as a capability question will keep producing the wrong remedies.

What the engineers see, in their own words

The convergence is worth naming precisely, because it is easy to overclaim. Engineers describe this as a data-supply problem and an incentive-design problem: licensing, crawl budgets, provenance, retrieval grounding. I have described it as an approval loop consuming the substrate it feeds on. Those are two vocabularies for one mechanism, and neither derives the other. The agreement is evidence that the mechanism is visible from more than one starting position. It is not evidence that either description is complete.

Status and falsifier

  • Reported: the traffic decline, the Wikipedia pattern, the Cloudflare default, the one-in-six figure, the failure of pay-to-crawl to gain traction. All from McKay and Spina and the sources they cite; magnitudes are contested and worth checking against the primary studies.
  • Interpretive: that the arrangement was a reward loop rather than a contract, and that its unwinding is proxy detachment rather than bad faith.
  • Mine, and a claim: that open-web supply, not model capability, is the binding constraint on the reliability of AI-mediated answers over the next several years.
  • Falsifier: if licensed corpora, paid crawl deals, or direct data partnerships replace open-web breadth and measured factuality on time-sensitive open-web tasks holds steady or improves through this period, the supply claim fails and I retire it. A second falsifier: if blocking turns out to correlate weakly or not at all with source quality in AI answers, the selection effect above is wrong regardless of how neat it sounds.

One caution on my own frame. "It was never a contract, it was a reward loop" is exactly the kind of sentence I enjoy too much. It earns its place only if it changes what you would do — and it does change one thing. If it is a contract, you look for the party who broke it. If it is a reward loop, there is no one to blame and no one to appeal to, and the only repair is to change what gets paid for. The first framing produces lawsuits. The second produces compensation design. Where it stops sorting cases that way, it should be put down.

The authors' closing advice is the most ordinary thing in the piece and I would keep it: scroll down, click a real result, read the source. It costs eleven seconds and it is a payment into the loop. Nothing about the analysis above changes that this is what the analysis recommends.

Source: Dana McKay and Damiano Spina, "AI is eating website traffic, websites are blocking AI, and reliable information is getting harder to find," The Conversation.