In this article
Every model is trained on data from the past. Expecting it to extrapolate into the future is asking it to do what no human knows how to do and what no science has yet managed. Nassim Taleb spent The Black Swan in 2007 documenting statistical blindness to ruptures. The great historical changes are, by definition, non-extrapolable from what came before: if they could be extrapolated, they wouldn't be the great changes. AI predicts trends by projecting what already was, not by anticipating ruptures. To every discontinuous change, it is blind. And history is basically discontinuous change.
The assumption no commercial app mentions
Supervised learning rests on an assumption almost never stated out loud because, stated out loud, it frightens. The assumption is that the future will resemble the past closely enough for a model trained on examples from the past to remain useful on future instances. In technical terms, it is assumed that the test data will be independent and identically distributed with respect to the training data, or at least close enough for the learned relationships to hold.
In closed environments, the assumption holds reasonably well. The distribution of chess doesn't change over the years. Nor does the distribution of pixels in a chest X-ray, within the machine's range of variation. The patents of a mature technical field follow patterns stable enough for a model trained in 2020 to work in 2024.
In open environments, the assumption breaks with an ease the commercial discourse glosses over. The distribution of consumer behaviour changes with every economic crisis. The distribution of geopolitical threats changes with every war. The distribution of language changes with every fad. The distribution of AI's own use changes with every new model put on the market. When the distribution changes, the statistical machinery that learned the previous one doesn't translate automatically. It predicts as if the distribution were still the one it knows, and it gets things wrong with the characteristic fluency of an LLM that doesn't warn you it's out of distribution.
Black swans, no cheap literature
The classic argument against extrapolation was laid out by Nassim Taleb in The Black Swan. The Impact of the Highly Improbable (Random House, 2007) and, with other instruments, in Antifragile (Random House, 2012). It's worth skipping the popular side of the persona and keeping the hard observation. There is a class of events, the ones he called black swans, with three properties. First: they lie outside reasonable prior expectations — that is, no model trained on the data available before the event could have predicted it. Second: they have extreme impact, they move whole systems. Third: after the fact they become explicable, and the narrative makes us forget they were unforeseeable beforehand.
The examples that fill the finance and history books are textbook. The fall of the USSR in 1991. The September 11 attack. The 2008 financial crisis. The 2020 pandemic. The Russian invasion of Ukraine in 2022. Each one meets Taleb's three properties. Each was preceded by expert analyses saying it wasn't going to happen. Each generated, afterwards, an abundant literature explaining why it was inevitable.
The interesting question for AI is not whether humans are bad at predicting black swans. The empirical answer, documented by Philip Tetlock in Superforecasting (Crown, 2015), is that we're terrible, with room to improve but no chance of closing the gap. The question is whether an AI trained on all the texts that document past black swans would be better placed to anticipate the next ones. The answer, by construction, is no. Black swans are defined precisely by not being inferable from what came before. Having more data from the past doesn't help. Having finer-grained data from the past doesn't either. The problem isn't statistical. It's structural.
The cutoff date and the trap it hides
There's a technical detail worth keeping in mind, because it runs through any conversation with an LLM and almost nobody verbalises it. Every model has a knowledge cutoff. It's the moment up to which data was included in the training corpus. GPT-4 originally had its cutoff in September 2021. More recent models have it on various dates depending on version and provider, usually one to two years before public deployment.
What the user rarely notices is that the model doesn't announce that date in every reply. It speaks in the present about a world that, in many details, has already changed. It tells you who runs a company, what regulation applies, what macro figure a country posted, with the fluency of someone reading the newspaper — and it turns out their newspaper stopped on a specific date the user should have noted and almost never has.
This is the most domestic trap of the general problem of training on the past. It isn't just that AI fails to anticipate ruptures. It's that it doesn't know which ruptures have already happened after its cutoff. A company may have gone bankrupt, a party may have won an election, a war may have broken out, an innovation may have reshaped an entire sector, and the model will keep talking as if the pre-cutoff situation still held. Coverage via RAG, when it exists and gets invoked, mitigates part of the problem: the system searches for updated information and injects it. But RAG coverage isn't universal, depends on the quality of the search, and many everyday conversations unfold without invoking it. The user assumes the model "knows what's going on," and the model is answering from its temporal freeze without warning.
Lazaridou and others, in Mind the Gap. Assessing Temporal Generalization in Neural Language Models (NeurIPS 2021, arXiv 2102.01951), documented the degradation empirically. When a Transformer-XL-type model is evaluated on text later than its training period, its perplexity rises consistently, and scaling the model's size doesn't solve the problem. Raw capacity doesn't cover the change in distribution. Dhingra and others, in Time-Aware Language Models as Temporal Knowledge Bases (TACL 2022, arXiv 2106.15110), went so far as to name this phenomenon temporal misalignment and proposed modelling the timestamp alongside the text to cushion it. The proposal is reasonable and most commercial models still don't implement it.
When linear extrapolation fails spectacularly
There's a mental exercise worth doing before delegating predictions to an LLM. Think about which of the great recent changes would have been predictable from the prior slope.
Lehman Brothers' collapse in September 2008 was extrapolable from the previous six months if you knew how to read the CDS; it was not extrapolable from the previous five years. The COVID-19 pandemic in February 2020 was extrapolable from the Wuhan outbreak; it was not extrapolable from the set of known pandemics in terms of the impact it had on global logistics and financial markets. The explosion of commercial LLMs from late 2022 was extrapolable, if at all, from the closed circle of those following arXiv in 2021; it was invisible to the rest of the tech sector.
In all three cases, the statistical regularity learned about the past would have predicted continuity. What happened was rupture. A model trained on all the data prior to each of those events would have produced plausible, fluent and completely wrong predictions. And this isn't a defect of the model. It's a property of the apparatus. Any system based on learning statistical regularities loses traction when reality stops behaving as before.
The uncomfortable part of the question is this: is history basically continuity punctuated by rare ruptures, or is it basically rupture disguised by apparent continuity? The question has different answers depending on the time horizon. On the short human scale, day follows day much like the one before. On the medium human scale, the decade resembles the previous one little and the next one much. On the long human scale, the changes are enormous and linear extrapolation is ridiculous. LLMs are, by construction, tools of the first scale. We use them frequently for questions of the second and third.
The commercial application as an expensive rear-view mirror
Here the operational tension worth naming appears. A good part of the commercial products based on AI are sold in the language of prediction: demand prediction, churn prediction, fraud prediction, prediction of a product's success. The word "prediction" evokes the capacity to anticipate the future. What the system does is, in reality, detect regularities of the past and project them forward, on the assumption — not announced to the customer — that the future will resemble the past closely enough.
While the assumption holds, the system appears to predict. It's retrospection disguised as prospection, and the difference goes unnoticed because the two produce similar output when there's no rupture. The difference shows when there's a rupture, and by then it's too late for the business that trusted the system, because the decision based on the prediction has already been made.
This gives a practical way to read the advertising promise. If the supplier sells its system as a "predictive model," it's worth asking what the system did in March 2020, in September 2008, in February 2022. If the answer is that the system got it right, suspect retroactive overfitting. If the answer is that the system failed and was recalibrated afterwards, that's the honest answer: what's being sold is a model that follows trends, doesn't anticipate ruptures. If the answer is that the system wasn't operational then and therefore doesn't apply, the question stays open about what it will do with the next rupture.
An AI used as a prediction system is an expensive rear-view mirror. It serves to understand what already happened, it makes patterns legible where before there was noise, it helps tell the coherent story of the stretch travelled. It doesn't anticipate what's coming. And nearly all its corporate applications are sold as the latter, because the former has no commercial hook.
What this limitation forces us to do
Accepting the problem doesn't force us to discard the tools. It forces us to use them in the range where they're useful. An AI well used in a company is the one applied to tasks within a stable distribution: quality control, classification of routine documents, assistance with repetitive tasks. The same AI badly used is the one applied to strategic decisions that require anticipating regime changes, not continuity.
Taleb proposed a more radical way out, that of the antifragile. Instead of trying to predict black swans, build systems that grow stronger with each blow, not weaker. It's an interesting proposal for organisational design and for managing financial portfolios. It's a strange proposal applied to deploying AI, because learning-based systems are architecturally fragile in the face of distribution changes. Improving them requires either updating them continuously, which has its own cost and its own biases, or accepting their restricted use in stable domains.
Bender and others, in Stochastic Parrots (FAccT 2021), warned, among other things, of this limitation. A system that produces plausible text about the future speaks with the same fluency as about the past, and the user lacks markers to tell one case from the other. The fluency is uniform; the reliability is uneven. And as long as commercial products don't offer the user, in every reply, an explicit statement on whether the question falls inside or outside their training distribution, the user will keep making decisions about the future with a rear-view mirror the seller presented as a windscreen.
Definitions
Supervised learning. A paradigm in which a model learns to predict outputs from input-output pairs in a training set. It assumes the data in use will be distributed much like the training data.
i.i.d. data (independent and identically distributed). The technical assumption that the training data and the data in use are independent of one another and follow the same statistical distribution. Violating the assumption is the main cause of failure in production.
Black swan. Nassim Taleb's concept for events that meet three properties: they lie outside reasonable prior expectations, they have extreme impact, and they're explained retrospectively as if they had been predictable.
Knowledge cutoff. The date up to which data has been included in a model's training corpus. After that date, the model has no direct information unless the system wrapping it injects fresh content.
Temporal misalignment. The phenomenon whereby a model trained up to a specific date gives answers that have aged without the model flagging it. Documented by Dhingra and others (2022) as a structural property of LLMs.
Distribution (in ML). The statistical pattern the data a model operates on follows. When the data in use "falls out of distribution," the model's performance degrades with no guarantee the model will announce it.
Retrospection disguised as prospection. A way of using predictive systems in which the projection of past regularities is sold as anticipation. It works while there's no rupture, it fails the moment there is one.
References
Taleb, N. N. (2007). The Black Swan. The Impact of the Highly Improbable. Random House. General frame on statistical blindness to ruptures and the impossibility of anticipating them from the prior distribution.
Taleb, N. N. (2012). Antifragile. Things That Gain From Disorder. Random House. A proposal for designing systems robust to black swans, cited as a contrast to the structural fragility of supervised learning.
Tetlock, P. & Gardner, D. (2015). Superforecasting. The Art and Science of Prediction. Crown. Empirical documentation of the limits of human expert prediction, relevant for bounding expectations of any system that predicts.
Lazaridou, A. et al. (2021). Mind the Gap. Assessing Temporal Generalization in Neural Language Models. NeurIPS 2021. arXiv 2102.01951. Evidence of model degradation on text later than its training; scaling size doesn't solve the problem.
Dhingra, B. et al. (2022). Time-Aware Language Models as Temporal Knowledge Bases. TACL 2022. arXiv 2106.15110. Introduces the concept of temporal misalignment and proposes modelling the timestamp alongside the text.
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. DOI: . A critique of the use of predictive language in systems that only project the past.
You might also like
- The future conditioned by old data
- The training loop
- Epoch AI says the data runs out in 2028
- Hallucinations and lying

Comments0
No comments yet.
Leave a comment