In this article
Shumailov and others published in Nature (2024) that training models on data generated by previous models produces "model collapse": progressive and irreversible loss of the tail of the original distribution. AI learns from humans. Humans increasingly generate text with AI. The next AI learns part of itself from a previous AI. It's a loop nobody planned and that's already running. Its consequences will be known after the fact. It's the biggest experiment we've ever run with no protocol.
The technical phrase almost nobody uses
"Recursive training on model-generated content." The phrase sounds academic and so it goes unnoticed. What it describes is this: a language model learns from a corpus. That corpus, today, contains human text and text generated by earlier models, with no operational distinction between the two. The next model, on its next training run, gets the mix. And we call the distillate of that mix "the advance of the state of the art." It isn't an obscure metaphor. It's the literal mechanics of the sector in 2026.
Shumailov and others measured it in AI models collapse when trained on recursively generated data (Nature 631, 2024). They replicated the loop in three different families: Gaussian mixture models, variational autoencoders and large language models. In all of them the result was the same: after several generations of training the successor model on the predecessor's output, the output distribution contracts. The rare tails go first. What's infrequent — the local idiom, the obsolete word, the culturally peripheral nuance — leaves first. The central stays and becomes more central. The statistical diversity collapses onto the median. They called this model collapse and underlined an uncomfortable property: the damage isn't aesthetic, it's irreversible. Models trained on contaminated data don't recover the lost tail, not even by adding more contaminated data afterwards.
The proportion no one wants to look at
Here the operational question appears: how much of current models' corpus is already AI output? Nobody publishes the figure. The big platforms don't either. But there are indirect indicators worth putting on the table. Elazar and others, in What's In My Big Data? (ICLR 2024), built an audit platform called WIMBD that analysed more than thirty-five terabytes of popular training corpora — C4, The Pile, RedPajama. What they found wasn't a clean proportion of synthetic versus human, because that proportion can't be measured directly once the content is mixed. They found massive contamination with benchmarks, high prevalence of duplicated content, low-quality content with no origin label, and personally identifiable information. The technical conclusion wasn't "the corpus is X per cent contaminated by AI." It was that there's no reliable way to know after the fact.
The second indicator is from the market. Originality.AI has published since 2023 a continuous tracking of Google search results. In its 2025 samplings, around 17–20% of the top-20 results contain AI-generated content. The figure has to be taken with the caution a company that sells AI detectors deserves: the incentive bias is there. What's significant isn't the exact percentage. It's that, whatever the real figure, it's already high and rising. And the crawler that feeds the next training corpus ingests that mix with no distinction, because there's no scalable protocol that lets you tell human from synthetic when the synthetic is well made.
How the tail is lost, said without jargon
The concept of the "tail of the distribution" tends to sound abstract. A concrete example helps. When a language model is trained on contemporary peninsular Spanish, there are mid-frequency words appearing thousands of times — casa, coche, trabajo — and there are rare words appearing few times — desbravar, otear, quincenal. The first form the body of the distribution; the second form the tail. When a G1 model generates text from its learned distribution, it tends to use the first with a frequency close to that of the original corpus and the second with a somewhat lower frequency, because generation samples toward the probable. If a G2 model is then trained on text generated by G1, the rare words appear even less in G2's corpus. After several generations, the tail disappears. The successor speaks with a narrower vocabulary than the predecessor, and the predecessor already spoke with a narrower vocabulary than the original human corpus.
The phenomenon scales beyond the lexicon. Rare syntactic structures, minority cultural formulations, localised geographic or historical references, subculture expressions, marked registers — anything infrequent in the original human corpus loses weight with each generation. What remains is a progressively more homogeneous, more global, more predictable cultural median. The model doesn't get stupider. It gets flatter.
The legal problem being cooked up
There's an ongoing debate in the Harvard Journal of Law & Technology about whether a new legal category is emerging, the right to uncontaminated data, understood as the right that there exist identifiable corpora of audited human production from before the generative era, preserved by public institutions, accessible for research and for training models that need to anchor their distribution to a verifiable human baseline. The proposal hasn't consolidated and the details vary between authors. The core is reasonable: if the training loop keeps accelerating and no identifiable human reserves remain, the technical viability of future generations of models depends on artificially preserving the human baseline.
The interesting part here is the cultural inversion. Until five years ago, the public conversation about data revolved around the protection of human privacy from systems. Now it's starting to revolve around the protection of systems from the flood of their own output. The subject and the predicate swapped sides and almost nobody named it. It's the sector equivalent of asking that the spring's purity be protected because the factory downstream needs clean water to produce the bottled water that goes back to the spring.
The suspicion of fluency
Bender, Gebru and others, in On the Dangers of Stochastic Parrots (FAccT 2021), warned beforehand that the dominant quality metric in LLMs — the fluency of the generated text — was a deceptive metric, because fluency is a property of syntactic-lexical performance that a model trained on massive corpora exhibits almost by construction, and that doesn't covary with truthfulness or with diversity. The argument is relevant here because, in the presence of model collapse, the first thing maintained is precisely fluency. The tail that disappears isn't the surface quality. It's the depth of richness.
This produces a paradoxical effect easy to mismeasure. The user reads a text generated by the latest version of the model, checks that it's coherent, grammatically impeccable, naturally readable, and concludes the model "has improved." That reading is correct on the surface plane and completely blind to the deep plane. A narrower model can be more fluent. A more diverse model can be rougher. The blind preference for fluency is a blind preference for the collapsed model.
What the engineer knows and the company plays down
The technical sector knows the problem. The big labs have invested in filters that try to detect synthetic content, in datasets with certified-human provenance markers, in contracts with content producers to secure a clean corpus. Crawford, in Atlas of AI (Yale UP, 2021), already then documented the hidden economy of annotation work that props up the quality datasets. That economy is growing. There are startups selling "pre-2022" human corpus as a high-end product, with the same logic by which museums sell collections of copper mined before the atomic tests of the forties, because that copper is the only one not contaminated with radioactive isotopes of the nuclear era and therefore serves to manufacture detectors needing a clean baseline.
The analogy isn't rhetorical. It's functionally exact. There's a background contamination distributed globally, irreversible in the short term, indistinguishable to the eye, that forces precision-instrument makers to seek reserves predating the contaminating event. The big AI labs are doing exactly the same with the human text corpus: identifying, sealing off and protecting reserves dated before the generative flood, not for the common good, but because without them you can't manufacture the next generation of models with the statistical diversity the next model needs in order not to collapse.
The experiment already running
The uncomfortable question isn't whether the loop happens. It happens. The question is at what speed and with what accumulated cultural consequences. And here there's an observation worth making without flourish. There is no experimental protocol. There's no control group. There's no point of return. There are no predefined stopping conditions. There's no institutional authority with the capacity to halt the recursive ingestion while it's measured. The experiment is conducted on the global linguistic population, with consequences that will manifest in the cognitive, expressive and aesthetic habits of the next generation of users.
It would be sensible, if this were science, to demand an IRB, an ethics review, a monitoring plan. It isn't. It's competitive commercial deployment, and the question of whether the cultural deep plane flattens or not gets answered at the speed of three financial quarters. The question will reach the regulations once it already has an empirical answer. That empirical answer will be known after the fact, when the effects are visible. And when they are, they'll already be incorporated into the very corpus from which one might reflect on them.
Will we be able to write about the loop, ten years from now, with words that don't come from the loop? It's a question with no grandiloquence. It's operational.
Definitions
Model collapse. A phenomenon documented by Shumailov and others (2024) whereby training generative models on the output of previous models produces progressive and irreversible loss of the tail of the original distribution. The rare samples disappear first; the output becomes more fluent and narrower.
Tail of the distribution. The part of a statistical distribution where the infrequent but significant events reside. In a language model, the tail holds the rare lexicon, the minority structures and the peripheral cultural references. Its loss is undetectable by fluency metrics.
Recursive training. The practice of training a model on a corpus that contains output from earlier models. When the cycle repeats across several generations without provenance filtering, it produces model collapse.
Data provenance. A verifiable mark of a content fragment's origin: human or synthetic, generated with which model, on what date, under what conditions. At web scale, there's no established protocol to apply provenance reliably and universally.
Human baseline. An identifiable corpus of audited human production, generally dated before the mass spread of generative models, kept as a statistical reference for training and evaluating future models.
References
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? Proceedings of FAccT 2021. An early critique of fluency as a deceptive quality metric and of the cultural implications of massive training.
Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. An analysis of the material and labour chain of model training, including the hidden economy of quality datasets.
Elazar, Y. et al. (2024). What's In My Big Data? ICLR 2024. arXiv: 2310.20707. The WIMBD platform for auditing training corpora; documents the impossibility of directly measuring the proportion of synthetic content once integrated.
Harvard Journal of Law & Technology Digest (2025). Model Collapse and the Right to Uncontaminated Human-Generated Data. Published March 2025. A recent legal discussion of the emergence of a category of right related to preserving an audited pre-generative human corpus.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N. & Anderson, R. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755–759. Empirical demonstration of model collapse in VAEs, GMMs and large language models, with analysis of the irreversibility of the distributional tail loss.
You might also like
- The degradation of knowledge
- The entropy of digital content
- The collapse of content's value
- The dataset as ideology

Comments0
No comments yet.
Leave a comment