Epoch AI Says the Data Runs Out in 2028

In this article

  1. The paper, what it says and what it doesn't
  2. The four exits from the dead end
  3. Why synthetic data doesn't save us, yet
  4. The new private gold
  5. What changes for you, the user who signs nothing
  6. The political question
  7. To go deeper
  8. You might also like

Definitions · References · Elsewhere

Today we're talking about the 2028 horizon. Epoch AI, a small and very serious organisation, has spent years publishing that the high-quality public data for training large language models will run out around that date. The prediction has been through every technical publication and no general-interest press. Why does it matter to us? Because when the public corpus runs out, the model with the advantage will be the one with access to a private corpus. And that turns every conversation of yours with a chatbot into a patrimonial decision without anyone having asked you to sign. My take is biased because I've been careless with my data and have regretted it. Form your own.

I've spent six months with Epoch AI open as a fixed tab. It's one of the few groups that publishes falsifiable predictions, with a public methodology, about the sector's trajectory. The conclusion we'll deal with today is known as Will we run out of data? and is probably the group's most-read work. It's been cited in dozens of academic papers. It's barely cited in the Spanish-language press. I wanted to fix that.

The paper, what it says and what it doesn't

The work is signed by Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn and Anson Ho, and the original version is from November 2022 (arXiv:2211.04325). What we call "the 2028 paper" is really an update published in 2024 in the Position track of ICML, Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data. Between the initial version and the update the figures changed, and the change says a lot about the problem.

The methodological idea is simple. The authors estimate, first, how much high-quality human text exists accessible online —discounting duplicates, spam, autogenerated text and low quality. Then they estimate at what rate that quantity grows. Then they estimate at what rate the appetite of the models grows —that is, how many tokens each generation consumes. And they project the crossing point.

In the 2022 version, the crossing fell near 2024 in a central scenario, and between 2026 and 2032 in optimistic and pessimistic scenarios. That set off the alarms. In the 2024 update the authors revised the estimate of available text upward —the fine filtering of the web had improved and let them exploit more data than were counted as usable. The estimate of high-quality text stock went from around 10 trillion tokens to around 50 trillion. And, with that revision, the crossing shifted to roughly 2028.

The figure of 300 trillion tokens that sometimes appears in headlines refers to the total stock of available public human text, including low quality. Frontier models don't use that whole set. They use the filtered, much smaller one. That's why the figure that matters is the 50 trillion, not the 300.

There are caveats to the forecast worth keeping in mind. The first is that "running out" doesn't mean zero. It means the ratio between data consumed by the run and new data produced each year unbalances to the point of making it unprofitable to keep scaling by the classic route. The second is that the prediction only concerns public textual data in major languages. There are alternative reservoirs —video, audio, code, science, minority languages— that the curve doesn't affect the same way.

Even with all the caveats, the central result is stable: if the current trajectory of scaling by raw data continues without structural changes, around 2028 there ceases to be enough high-quality public text to keep feeding it.

The four exits from the dead end

When you approach a bottleneck, there are four paths. It's worth seeing them separately because each defines a different group of actors in the sector.

The first is training with synthetic data generated by other models. Attractive in theory —if a big model can produce text, we can use that text to train the next. There's a growing technical literature defending it, especially for data curated with outcome verification, like the ones DeepSeek used in training R1. But the naive version of this strategy has a documented problem, which I return to in the next section.

The second is training with private data of human quality. That is, with text produced by humans but not accessible on the public web. The internal archives of the big publishers. The conversations on closed platforms. The chats with AI assistants. Corporate emails. The recording of calls. Academic writing. This is where the sector has been betting aggressively since 2023. We see it in OpenAI's licence agreements with News Corp, Axel Springer, the Financial Times, Le Monde and other publishers. We see it in Google's deal with Reddit signed in February 2024 for an annual figure on the order of 60 million dollars according to the leaks gathered by Reuters. We see it in Meta using public Instagram and Facebook content as a training corpus since 2024. We see it in Microsoft accessing LinkedIn through corporate structure.

The third is expanding toward modalities other than text. Video, audio, image, robotic sensors. The corpus available in YouTube video and in podcast audio is orders of magnitude larger than that of text, once digested by multimodal models. This is one of the strongest lines of investment for Google and OpenAI since 2024. It doesn't solve the bottleneck for purely textual tasks, but it opens new dimensions.

The fourth is changing the training paradigm. Reducing the dependence on raw volume and increasing the dependence on reasoning. Extended-reasoning models like OpenAI's o1, DeepSeek's R1, or the Claude generations with extended thinking, demonstrate that a model can improve on hard tasks by dedicating more compute to inference, not more data to training. This line is, probably, where the most interesting things are happening today.

Those four exits aren't mutually exclusive. Any frontier lab will combine all four. But the combination each one makes decides, in good part, where it competes. And only two of the four are really within anyone's reach: the other two require size and capital very few have.

Why synthetic data doesn't save us, yet

In July 2024, the journal Nature published a work titled AI models collapse when trained on recursively generated data, signed by Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson and Yarin Gal (Nature 631:755-759, 2024). The paper documented with rigorous experiments what the sector suspected but hadn't measured. When a model is trained on the output of earlier models, generation after generation, it suffers what the authors called model collapse: it loses the tail of the original statistical distribution, becomes more uniform, worse calibrated, and ends up reproducing only the most frequent patterns.

The reason is statistical before it's philosophical. Each model generation learns the rare events of the original corpus a bit badly, and when it generates text it underproduces them. The next generation, trained on that generated text, sees them even less. The third generation loses them. The lexical, semantic and stylistic diversity of the corpus compresses toward the mode. The model becomes, in technical terms, peakier —more pointed and less broad.

The paper received legitimate rejoinders. Some researchers have shown that when synthetic data is filtered carefully and mixed with human data in controlled proportions, the degradation can be mitigated. DeepSeek and other Chinese labs demonstrate that it's possible to train with partly synthetic corpora without collapsing, provided the synthetic is verified against external ground truth. But the naive version of "we train each model with what the previous one generated" is closed.

This leaves synthetic data as a complementary tool, not a substitute. The easy exit from the bottleneck is no exit.

The new private gold

Once the two cheap exits close —limited public data and non-trivial synthetic data— what remains as a competitive-advantage resource is private data of human quality. And this is where the sector's economics becomes uncomfortable to read.

The strategic move of the big hyperscalers since 2023 is best understood read in terms of corpus acquisition. Microsoft acquired LinkedIn in 2016 for 26 billion dollars and keeps preferential access to the stream of professional text millions of professionals write each day. Google signed with Reddit the 2024 deal that gives it access to the conversations of the most relevant general forum in the world in the English language. Meta uses public Instagram and Facebook content as a corpus since 2024, and has officially confirmed this use, generating regulatory friction in Europe besides. Amazon accumulates interaction data from Alexa and transactions on its marketplace. Apple announced in 2024 agreements for training on news publishers' data. OpenAI keeps a growing series of licences signed with big publishers that, according to Press Gazette's public counts, already add up to multi-year commitments above a billion dollars.

This isn't conspiracy. It's visible strategy. The operational question is: where does the private corpus come from for someone who isn't one of these six or seven? The short answer: from the chat you keep with their model.

What changes for you, the user who signs nothing

When you open an account on an AI service and start conversing, you generate a corpus. Each service's privacy policy decides what's done with that corpus. Some —Anthropic Claude on its paid API, Microsoft Copilot Enterprise— promise not to use it for training. Others —the free version of many services— do use it, except for an explicit opt-out the user rarely activates because it's buried in settings.

That means every time a user formulates an interesting question, pastes a fragment of a personal document, drafts an email with an assistant, or talks about their life with a chatbot, they leave training gold for the model's next generation. The quality of that gold depends on the user. A doctor explaining a clinical case, a lawyer summarising a file, a programmer debugging a specific system, a writer refining a paragraph: each contributes a kind of data that in the public set is scarce and in the private one abundant.

The consequence is visible two years out. The company that gathers the highest-quality human private corpus between 2024 and 2028 will have a substantial advantage when the public corpus saturates. The public models will talk alike; the models trained with the chats of millions of professionals will be the ones that know how to solve the concrete problems of each user's trade.

This is neither a call to panic nor an invitation to abandon chatbots. It's an invitation to read the service's privacy policy before putting sensitive information there, and to understand that the model that seems free is charging in a currency called your text.

The political question

This is personal opinion but it rests on the economics I've been describing. The predicted exhaustion of the public corpus in 2028 turns private data into a national strategic asset. And that puts before regulators a question that's only beginning to be discussed: who owns the corpus generated by citizens' use of AI services?

The answer today is: the service operator, unless the law says otherwise. In Europe, the General Data Protection Regulation (GDPR) recognises individual rights over personal data, but the regime for non-personal data generated during interaction with an AI system is fuzzy, and the big platforms operate in that fuzzy zone. The European AI Regulation (Regulation EU 2024/1689) covers obligations of transparency and traceability, but doesn't settle the patrimonial question.

In the United States there isn't even a federal regulation analogous to the GDPR, and data generated in chats is, by default, the platform's operational property. In China, the corpus is a state asset and is governed by national-security regulation. Three radically different regimes for the same raw material.

What I hold is that this regulatory difference is going to decide, over the next five years, where the raw material of the next training cycle accumulates. And, by extension, where the capital concentrates. Europe may end up the continent with the best individual protection and the worst capacity to retain corpus within its borders. The United States may end up capturing the global corpus on its platforms. China may end up with a closed national corpus but no access to the global one. Those three scenarios are being drawn now, and the concrete figure that decides when the bottleneck is reached is the one Epoch has been publishing since 2022.

The most recent update of the work by Villalobos, Sevilla and collaborators, Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data, presented at ICML 2024 (PMLR 235, July 2024), puts the crossing at roughly 2028 under the central scenario, assuming constant training efficiency. If that efficiency improves —and the DeepSeek case suggests it does, see my 0055— the date can be pushed a year or two. If efficiency doesn't improve, it arrives sooner.

Definitions

Training corpus: the set of data —in this case, text— used to fit a language model's parameters. The quality and diversity of the corpus largely determine the quality of the resulting model.

Token: the minimal unit of text processed by a language model. In Spanish, a token usually equals about 0.75 characters on average, though it depends on the tokeniser. The 50 trillion tokens equal roughly the text contained in hundreds of millions of books.

Model collapse: the progressive degradation of a model when trained recursively on the output of earlier models. The statistical diversity of the original corpus is lost and the model becomes more uniform and worse calibrated. Documented in Shumailov et al., Nature 2024.

Compute-optimal training: a training regime that optimises the relation between model parameters and corpus data for a given compute budget. The Chinchilla scaling laws (Hoffmann et al., DeepMind, 2022) are the most-used formulation.

References

Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, Anson Ho, Will we run out of data? Limits of LLM scaling based on human-generated data (arXiv:2211.04325, 2022; revised version in Position: Will We Run Out of Data?, ICML 2024). The original paper and the update with the 2028 figure.

Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, AI models collapse when trained on recursively generated data (Nature 631:755-759, July 2024, doi:10.1038/s41586-024-07566-y). The reference paper on model collapse.

Epoch AI, Will we run out of data to train large language models? (epoch.ai, blog post and ongoing update). An accessible summary of the paper and the methodology.

Reuters, coverage of the Reddit-Google deal of February 2024. Source of the annual licence figure.

Press Gazette, aggregate coverage of OpenAI's licence agreements with publishers (2023-2025). Combined figures of multi-year commitments.

Hoffmann et al., Training Compute-Optimal Large Language Models (DeepMind, arXiv:2203.15556, 2022). The Chinchilla scaling laws that frame the debate over data vs parameters.

Stanford HAI, Artificial Intelligence Index Report 2026 (April 2025). The chapter on availability and use of training data.

To go deeper

Vili Lehdonvirta, Cloud Empires (MIT Press, 2022). Why the platforms that control the corpus of citizen use are building a power that exceeds that of states.

Daron Acemoglu & Simon Johnson, Power and Progress (PublicAffairs, 2023). A historical framework for understanding how the benefit of a technology is distributed when a key resource concentrates.

Ali Borji, A Note on Shumailov et al. (2024) (arXiv:2410.12954, October 2024). A nuanced rejoinder to the Nature paper; useful for not accepting model collapse as dogma.

Yves Citton, L'Économie de l'attention (La Découverte, 2014). A conceptual antecedent of the shift from attention to data as the dominant commodity.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment