Generative search doesn’t just route around publishers. In this paper, it is modeled as a force that can drain the crawlable web itself by capturing value from crawled content without sending traffic back to the source. That matters because the web is treated here not as a fixed indexable input, but as a renewable commons whose stock depends on whether publishers can keep funding new content.

That shift in framing is the paper’s main contribution. Once you treat crawlable content as a common-pool resource, the familiar questions about answer quality and click loss become sustainability questions: how much extraction can the system absorb before volume, quality, and lifetime all start falling together?


The crawlable web is not a fixed input — it is a renewable commons

The paper’s starting point is simple: generative search engines answer queries directly from crawled web content, and the value they capture often does not return to the source. The authors call that capture “extraction,” and they treat it as a diversion of the traffic that finances content production.

That leads to a different model of the web. Instead of assuming the crawlable corpus is just sitting there as a stable input to search, the paper formalizes it as a common-pool resource: the crawlable commons. Like any commons, it can be overused. And here, overuse does not mean fewer documents in a static archive. It means the stock stops renewing itself.

The commons is described by three quantities: volume, average quality, and lifetime. Those are not three separate problems. They move together once extraction starts breaking the publisher feedback loop.


Three publisher responses create three kinds of erosion

The paper argues that extraction degrades the crawlable commons through three channels at once: dilution, depletion, and decay. That is the core mechanism, and it is more interesting than a simple traffic-loss story because the damage shows up in composition, investment, and durability simultaneously.

Dilution happens when publishers opt out. The crawlable corpus then changes in composition because some of what used to be available is no longer crawlable. Depletion happens when a publisher stays but earns less, so lower revenue reduces the production of new content. Decay happens when publishers shift toward perishable content, because durable pages are the ones the engine extracts best, which makes shorter-lived content relatively more attractive.

The point is that extraction does not just shrink the corpus in one dimension. It makes the remaining corpus worse to crawl, less worth renewing, and more fragile over time.


Cross the erosion threshold, and the commons goes extinct

The model has a tipping point. Once extraction rises above the erosion threshold, the crawlable corpus eventually goes extinct. That is the paper’s sharpest result, because it turns extraction from a gradual nuisance into a sustainability boundary.

The threshold also separates different kinds of engine behavior. A myopic GSE can push the system past it, while a long-run oriented GSE stays below it. So the difference between short-horizon and long-horizon optimization is not cosmetic here; it decides whether the commons survives.

The paper also proves, under a concavity condition on the steady-state value of the commons, that the symmetric equilibrium extraction rate is nondecreasing in the number of engines and converges to the threshold. In plain language: competition does not naturally stabilize the system. It can push extraction upward until it sits right on the edge of extinction.


The welfare result is stricter than the competitive one

The socially optimal extraction rate lies strictly below the erosion threshold, and it is no higher than the single engine’s sustainable optimum. That gives the paper a clean policy implication: whatever the market does, the welfare-maximizing choice leaves a buffer below the point where renewal collapses.

The result is intentionally conservative in a few ways. Several of the paper’s results restrict attention to stationary extraction rates. The competition theorem also relies on a concavity condition that is not implied by the first three assumptions, and it assumes equal splitting of queries rather than market shares. The welfare result is presented under the assumption most favorable to extraction: for a given corpus, users prefer more extraction.

Even with those pro-extraction assumptions, the optimal answer is still restraint.


What to do about it

If you are thinking about crawler policy, publisher monetization, or AI-search strategy, the useful move is to stop treating extraction as a neutral efficiency gain. This paper frames it as a sustainability externality. That means the relevant question is not just whether extraction improves answer delivery, but whether it keeps publisher incentives strong enough for the corpus to renew itself.

The practical implication is straightforward: policies and product strategies should be judged against an erosion threshold, not against short-run retrieval gains alone. If extraction rises too far, the system does not merely become less fair to publishers — it becomes less able to produce the content future search depends on.

Key Takeaway

The paper’s commons dynamics link extraction to three simultaneous degradation channels (dilution, depletion, decay), with a tipping point for system survival.

Treat generative search extraction as a sustainability externality: once extraction is high enough, the web stops renewing itself, so optimal policy/strategy must explicitly constrain extraction below the erosion threshold.

Source

Sylvain Peyronnet (2026). When Search Eats the Web: A Model of Corpus Erosion under Generative Extraction. arXiv:2608.15896