The entropy of digital content. All the noise within reach

In this article

  1. The founding promise and its operational flip side
  2. The number nobody publishes with precision
  3. The academic measure of the decline
  4. The "dead internet theory" as a partial description
  5. Synthetic SEO, the economy of cheap abundance
  6. The consequence the user's own metric does reflect
  7. The promise, reversed
  8. You might also like

Definitions · References · Elsewhere

An analysis by Graphite estimated that in November 2024 the amount of new AI-generated articles surpassed, for the first time, that of human-written articles on the web, putting the proportion around 50%. The figure is disputed and measures long articles classified by an automatic detector, not "all the text on the internet," but the direction of the trend leaves little doubt. There's more and more content and it adds less and less. The signal dilutes into noise — automatic generation, recycled content, duplicated posts, synthetic SEO. Searching for something specific on the internet in 2026 is already a quasi-archaeological act. The promise of "all knowledge within reach" is being fulfilled in reverse: all the noise within reach.

The founding promise and its operational flip side

In the late nineties the web sold a concrete promise: all the world's knowledge, accessible from anywhere, free or nearly so. It was a strong promise and, for a good decade, it held up better than seemed reasonable. Wikipedia, the academic repositories, the technical forums, the personal blogs, the great digitised libraries. Searching for an answer to a concrete question worked. The search engine returned results worth reading.

The promise is still printed on the marketing pages. What has changed is the practice. In 2026, searching for a concrete answer to a concrete question is an exercise that demands training. The user who remembers what searching was like in 2010 feels the friction. The user who knows only the present takes it as a normal condition: it's reasonable that the first screens of results are SEO farms, AI-generated content, unsigned listicles, aggregators of aggregators, and that the useful content appears, if it appears, hidden on page three or after adding three operators to the query.

This isn't nostalgia. It's a measurable change.

The number nobody publishes with precision

How much of the text published on the open web is already generative-AI output is a disputed figure and worth treating with the caution it deserves. The estimates that circulate have source bias: companies that sell detectors have an incentive to publish high figures, companies that produce models have an incentive not to publish them. There's no independent census with auditable methodology. The most cited datum, Graphite's, measures a bounded universe — long web articles classified by an automatic detector — and not "all the new text on the internet"; taking that figure as if it measured the web as a whole would stretch it beyond what its author claims. The analysis itself also qualifies that those AI-generated articles appear far less in Google's top results than their publication volume suggests.

The figures that are well documented refer to adjacent phenomena. NewsGuard, in its continuous tracking of sites identified as Unreliable AI-Generated News (UAIN), had catalogued more than a thousand sites whose production model is massive automated generation, a figure its own tracking centre describes as steadily growing. The figure is journalistic-commercial and worth taking as indicative. The qualitative sign, however, leaves little doubt: there's a growing, not small, volume of web production whose economic purpose is to occupy indexed space without giving a human reader anything that justifies the click.

The second sign is the change in user behaviour. It's a widely reported phenomenon, though without a single hard figure quantifying it: more and more people add the word "reddit" at the end of the Google query to dodge the first algorithmic results and reach forums with human answers. It's a patch, an evasion manoeuvre whose existence confirms what the figures try to measure.

The academic measure of the decline

Bevendorff and others, in Is Google Getting Worse? A Longitudinal Investigation of SEO Spam in Search Engines (ECIR 2024), built a longitudinal study of the SERPs (Search Engine Result Pages) of Google, Bing and DuckDuckGo in the genre of product reviews. What they found backs the user's intuition with data. A minority fraction of product reviews on the open web use affiliate marketing, but most of the search results across the three engines are from pages with affiliate marketing. The proportion has shifted: what is minority in production on the open web becomes majority in visibility. The study concludes by observing that the line between content and farm has become progressively blurrier, and that the situation "will surely get worse in light of generative AI."

The authors aren't polemicists. They're information-retrieval researchers publishing at a peer-reviewed conference. What they say without rhetoric is what the user notices without being able to put into words: the engine whose only function was to order the corpus by relevance has become porous to a class of content whose purpose is to manipulate the ordering by relevance, and the producers of that class of content have, in 2026, generation tools that let them produce at zero cost what once had a high cost.

The "dead internet theory" as a partial description

There was a meme, around 2021, claiming that most of the internet was already dead: that the vast majority of content was bot-generated, that the interactions were synthetic, that the surviving humans were navigating an already-dead space without knowing it. It was called the dead internet theory. It initially circulated as a conspiracy. The literal claim — that "everything" is a bot — was false. The directional claim — that a growing proportion of production and interaction is no longer human — has dressed itself in academic respectability as generative AI broadened its deployment.

Yoshija Walter, in Baudrillard and the Dead Internet Theory. Revisiting Baudrillard's (dis)trust in Artificial Intelligence (Philosophy & Technology, Springer, 2025), treats the concept as a philosophical problem rather than an empirical thesis, and it's worth reading it in that register. What matters for a serious discussion isn't whether the theory is literally true. It's whether its descriptive core — that the channel is losing its condition as a space of majority human communication — captures something of the medium's state. The answer, in light of the partial data available, is that it captures something. The proportion of human-to-human interaction without the mediation of systems that already produce and filter content has shrunk. Not to zero. But it isn't what it was.

Synthetic SEO, the economy of cheap abundance

It's worth naming the economic mechanism without moralising. Producing content costs money. A well-written article, with research and verification, takes hours. An AI-generated article on the same topic, decent enough to index, costs cents. The cost structure of someone producing quality content and that of someone producing synthetic SEO aren't comparable. And the search engine, as long as there's no robust provenance filter, ranks on the basis of signals — backlinks, freshness, keywords, time on page — that the synthetic producer knows how to manipulate as effectively as the human producer, at a much lower cost.

The result is predictable. The curve of content published grows exponentially. The curve of content cared for by a human author grows at human speed. The proportion between the two curves shifts year after year. And the search engine, which operates on the union of both, sees less and less signal per sample. It's Shannon entropy carried into the public informational channel.

Here it's worth avoiding the dramatic tone. Quality content still exists. Cared-for sites are still produced. The user who knows what they're looking for can find what they're looking for if they're willing to pay the cost of searching. The difference is that the cost of searching has gone up. And that, in economic terms, means information has stopped being free in the sense the founding promise sold. It's free to pay the author, it isn't free of the cost of searching. Bender, Gebru, McMillan-Major and Shmitchell warned in On the Dangers of Stochastic Parrots (2021) of exactly this scenario; Nicholas Carr had been arguing since The Shallows (2010) that web mediation alters the way you read and think, not just the available inventory; Kate Crawford recalled in Atlas of AI (2021) that the system producing and filtering that content is an industry with an economy of its own, not a neutral academic repository. All three warnings predate the massive generative deployment. All three are being borne out now.

The consequence the user's own metric does reflect

There's a datum worth attending to because it's the user's own. The accelerated adoption of conversational assistants — ChatGPT, Claude, Gemini, Perplexity — as a substitute for the classic search engine isn't a fad. It's a rational response to the deterioration of traditional search. The user who no longer quickly finds what they want on Google migrates to a system that gives them a synthesised answer on one screen, even knowing the answer may contain errors. They're paying the error in exchange for not paying the cost of searching.

It's a partially rational operation. The conversational assistant summarises over a corpus contaminated by the same problem. When it answers, it mixes hits about the mode with plausible errors about the tail. The answer arrives fast and, in many cases, sufficient. In others, false. And since the fluency is uniform and the reliability is uneven, the user has no easy markers to tell one case from the other.

The conversational interface solves the symptom without solving the disease. It makes the noise habitable. And in making it habitable, it legitimises the persistence of the noise, because it shifts the demand from the previous link — the functional search engine — to a later link content to synthesise what there is, without paying the cost of improving the source.

The promise, reversed

The phrase "all knowledge within reach" had the user as its subject and availability as its verb. The verb was fulfilled. The subject shifted. What is within reach of anyone with a connection is, today, all production: the cared-for, the cheap, the false, the honest, the generated, the copied, the translated, the repeated. Reach doesn't discriminate. Discrimination is a human operation, and the human cost of discriminating has risen at the same time as reach was being fulfilled.

Is this a failure of the promise or its literal fulfilment with unannounced side effects? Both readings are legitimate. What isn't legitimate is to keep selling it with the same face as in 1998. The phrase "all knowledge within reach" and the phrase "all the noise within reach" describe, already in 2026, the same object. The difference between the two isn't in the web. It's in the reader who crosses it.

Definitions

Shannon entropy. The formal measure of the informational content of a distribution. When a source emits more and more signals with less and less average surprise, its useful entropy per unit emitted decreases.

SERP (Search Engine Result Page). The results page a search engine returns to a query. Studies of search quality measure the proportion of relevant, spam or affiliate content in the top positions.

SEO (Search Engine Optimization). The set of practices for positioning content on the first pages of search engines. When those practices are applied to content whose aim is to index more than to inform, it's called SEO spam or a content farm.

UAIN (Unreliable AI-Generated News site). A category used by NewsGuard for websites whose production model is massive automatic generation with no verifiable editorial standards.

Dead Internet Theory. An informal thesis that a growing proportion of content and interaction on the web is generated or mediated by automatic systems rather than humans. Its literal version is false; its descriptive core points to a real quantitative change.

References

Graphite. More Articles Are Now Created by AI Than Humans. An analysis placing November 2024 as the point when the volume of new AI-generated web articles surpassed that of human-written ones — around 50% over a sample of long articles classified by an automatic detector — while noting that those articles appear little in Google's top results.

Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. An early warning on the cultural effects of the mass deployment of text generators.

Bevendorff, J., Wiegmann, M., Potthast, M. & Stein, B. (2024). Is Google Getting Worse? A Longitudinal Investigation of SEO Spam in Search Engines. ECIR 2024. An empirical study of the longitudinal quality of Google, Bing and DuckDuckGo in the product-review genre.

Carr, N. (2010). The Shallows. What the Internet Is Doing to Our Brains. W. W. Norton. A historical frame on the cognitive consequences of web mediation, predating the generative deployment but relevant for reading its continuation.

Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. A general reference on the economics and materiality of the systems that produce and filter content.

NewsGuard. AI Tracking Center. Tracking AI-enabled Misinformation. Continuous tracking of sites catalogued as UAIN, with growth documented by its own tracking centre.

Walter, Y. (2025). Baudrillard and the Dead Internet Theory. Revisiting Baudrillard's (dis)trust in Artificial Intelligence. Philosophy & Technology, 38:54, Springer. A philosophical discussion of the dead internet theory as a problem about the condition of the channel, beyond its literal reading.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment