Probabilistic truth. School teaches the right answer, the world asks for a degree of confidence

In this article

  1. The exam as a model of knowledge
  2. A tournament that measured who gets it right
  3. Why it doesn't reach the classroom
  4. The machine that asserts without knowing
  5. What gets decided in the voting booth

Definitions · Referencias · Para profundizar · También te interesa · En otros sitios

I was raised to find the answer. A single one, marked in red pen if I got it wrong. It took me years to realize that almost no adult decision that matters has that shape: the doctor doesn't know whether it's cancer, they estimate; the analyst doesn't know whether there'll be a recession, they weigh. Generative AI arrives disguised as oracle and worsens the misunderstanding, because it returns not truth but a well-calibrated bet dressed up as a confident statement. Whoever was trained to read in binary receives it as certainty; whoever learned to think in degrees receives it as a hypothesis to test. And thinking in degrees can be trained, even if school rarely tries.

The exam as a model of knowledge

I remember the format before I remember the content. The test, the single-solution problem, the dictation, the formula you apply that returns the number the teacher already had written down. All that scaffolding rewarded one thing, hitting the right answer to a closed question, and beneath it pulsed a silent idea about the nature of knowledge: the answer exists, out there, complete, and the student's task is to find it.

For much of the twentieth century the model was no absurdity. The knowledge base school transmitted was stable, and school had almost a monopoly on transmitting it. Two times two is four. The Battle of Hastings was in 1066. The Ebro empties into the Mediterranean. On that terrain binary is no defect but the exact tool, and demanding probabilistic nuance from someone learning basic geography would be pedantry.

The problem appears where knowledge demands judgment under uncertainty, and it turns out that territory is, as it happens, where the adult has something at stake. Medical diagnosis, economic forecast, risk assessment, legal interpretation, political decision: none of those fields admits a single right answer, all admit more or less probable answers. School trained us to operate as if life were arithmetic when life, in almost everything that counts, looks far more like a diagnosis.

A tournament that measured who gets it right

Between 2011 and 2015, Philip Tetlock and Barbara Mellers ran, from the University of Pennsylvania, the Good Judgment Project, one of the teams that competed in a forecasting tournament funded by the Intelligence Advanced Research Projects Activity (IARPA, the research agency of the US intelligence apparatus). The starting question was empirical and, for once, measurable: what distinguishes someone who forecasts geopolitical and economic events well from someone who forecasts badly? Thousands of volunteers, four years, hundreds of questions with later verifiable resolution.

The uncomfortable finding, gathered in Superforecasting, the book Tetlock co-wrote with Dan Gardner, was that the tournament's best forecasters —the ones the project christened superforecasters— beat, in calibration, professional intelligence analysts with access to classified information by more than thirty percent. They didn't win by IQ, which they had medium-high and not extreme, nor by a specialized training most of them lacked. They won by habits, and the habits hid no mystery.

Where almost everyone uses fuzzy categories, the superforecaster thinks in fine percentages: they say "sixty percent" where another would say "likely" and leave it at that. When new information arrives they don't jump from one certainty to its opposite, they adjust their estimate in small increments. They suspect their own hypothesis and go looking for what contradicts it before what confirms it. They break the coarse question into small questions they can attack separately. And they inhabit uncertainty without anguish, because for them "I don't know for sure" is a sensible starting point and not a confession of weakness.

The decisive piece for this matter is the last one, easy to overlook: all of that can be trained. A training module read in about an hour was enough, within the tournament itself, for the trained participants to beat the control group consistently across the four years. It wasn't a one-off improvement that evaporated within a month. If an hour of reading moves the needle for years, one might ask what a whole subject would do.

Why it doesn't reach the classroom

The explanation is not single, and it pays to distrust anyone who offers it neat. There are several causes, and they hold each other up.

There's inertia, to start. The education system was designed for another era and turns slowly; the curricula, the textbooks and teacher training take decades to change course, and nobody's in a hurry to move machinery that appears to work. There's also the comfort of grading: the binary exam grades itself almost at a glance, while a calibrated exam —where the student distributes probability among several options and is scored by how well-adjusted it was— turns out laborious and odd, even though international tests like PISA or TIMSS have incorporated the idea partially over the last decade. Add a curriculum already saturated, into which fitting reasoning under uncertainty forces something else out, and ministries prefer to accumulate content rather than replace it. And beneath it all sits the teaching staff themselves, who will hardly teach a framework their pedagogical training barely grazed.

Someone has been pointing at exactly this hole for thirty years. Gerd Gigerenzer, for years director of the Max Planck Institute for Human Development in Berlin and founder of the Harding Center for Risk Literacy, holds that the curriculum should include risk literacy. In Risk Savvy (Viking, 2014) he documents that doctors, judges, journalists, teachers and politicians make systematic errors interpreting elementary statistics, the same ones the average student reproduces because they imitated the binary mode they were taught.

His best-known example is the mammogram. Asked to interpret a positive screening, the bulk of doctors drastically overestimate the real probability that the woman has cancer: in controlled surveys they place it well above the true value, which sits below ten percent. It's not a calculation failure, it's a framing failure. When the same information is presented in natural frequencies —"ten out of a hundred" instead of a conditioned percentage— the error collapses. That, exactly that, is what school could teach in an afternoon and doesn't teach in twelve years.

The machine that asserts without knowing

It pays to look inside what a language model does, because the commercial jargon tends to hide it. Each word it generates (each token, the minimal unit of text it produces one at a time) is a sample from a distribution over the vocabulary, conditioned by everything before it. What the user reads is not a truth retrieved from a store; it's one of many plausible trajectories, the one that came out this time.

There's the trap. The fluency of the sentence covers up that the sentence is a bet. The model asserts "Madrid is the capital of Spain" with the same textual assurance with which it asserts "the main cause of inflation in 2024 was monetary policy," and yet the first has an almost unit probability and the second is debatable and depends on who you ask. Whoever doesn't think in degrees receives them identical, flat, equally certain, and acts accordingly.

From there a crack opens between two ways of using the tool. Without a probabilistic framework, the user treats the model as an oracle and limits themselves to asking, receiving and applying. With a probabilistic framework, they treat it as one more source that demands triangulation, and then they ask, receive, cross-check and decide. That the difference in benefit between the two profiles is large is suggested by the research on AI in skilled work: the experiment by Dell'Acqua and others on the "jagged technological frontier" showed that the same tool improves some tasks and worsens others depending on how it's used. The specific reading of that result in terms of probabilistic thinking is mine, not the study's.

It's not that AI invented the need to think in probabilities; what it did was raise the price of not doing so. Before, the citizen without that framework could survive by delegating to recognizable authorities, the doctor, the notary, the journalist. Now delegation gets complicated, because authority itself arrives mediated by probabilistic models nobody explains to them and that aren't transparent even to whoever sells them.

What gets decided in the voting booth

Here I come in, and I flag it: what follows is my opinion, though backed by what the literature on statistical literacy has been showing. A citizenry without a probabilistic framework is a manipulable citizenry. It swings between unfounded panic and false calm, can't read an evidence-based policy and, worst of all, doesn't realize it can't.

The recent examples pile up, and it pays to look at them closely because each fails in a different way. During the pandemic, the public debate on the effectiveness of measures, on mortality rates or on trial results rested on a poor statistical literacy: percentages were read as certainties, confidence intervals were ignored and a point forecast was cheerfully confused with the prediction of an accomplished fact. With climate something similar happens but in reverse. The IPCC reports use a calibrated and explicit scale —"very likely" equals a probability of 90 to 100 percent, "likely" 66 to 100— and the media translation grinds it down to a flat yes or no. And the old bias Kahneman and Tversky documented in 1974 persists, the one that inflates the fear of the vivid and rare, the attack, over the everyday and deadly, the car or the heart, steering public spending out of proportion to the real risk.

Democracy presupposes an electorate able to evaluate evidence, the evidence arrives in probabilistic format and school trains people to read it in binary. The mismatch is no anecdote for pedagogues.

The latest PISA data give the dimension of the hole. In PISA 2022, published by the OECD in December 2023, the OECD-country average in mathematics fell to 472 points, almost fifteen below 2018, the largest drop recorded between two consecutive editions. Within mathematical competence, the subdimension measuring reasoning with uncertainty and data remains among the weakest, and leaves a good part of students below the threshold at which an elementary statistical distribution gets interpreted. The binary learned in the classroom reappears intact, years later, in the hand marking the ballot.

Definitions

Calibration. A property of a forecast whose declared probability matches the real frequency of hits. Whoever says "seventy percent" calibrates well if, in the long run, they're right seventy out of a hundred times they say it.

Superforecaster. A forecaster identified in the Good Judgment Project for their sustainedly above-average calibration, capable of updating their estimates gradually in the face of new information.

Risk literacy. The ability, in Gigerenzer's terms, to interpret probabilities, frequencies and distributions applied to everyday-life decisions.

Natural frequencies. A way of presenting probabilistic information as "ten out of a hundred" instead of a conditioned percentage; it eases Bayesian reasoning and reduces interpretation errors.

Token. The minimal unit of text, a word or a fragment of a word, that a language model generates one at a time, sampling it from a probability distribution over the vocabulary.

Referencias

Philip Tetlock and Dan Gardner — Superforecasting: The Art and Science of Prediction (Crown, 2015). The article's main source for the habits of superforecasters and the results of the Good Judgment Project.

Good Judgment Project / IARPA. Geopolitical forecasting tournament (2011-2015) run by Philip Tetlock and Barbara Mellers at the University of Pennsylvania, basis for the data on accuracy versus intelligence analysts and on the effect of the training module. Verified at Good Judgment Inc. (goodjudgment.com/about) and the project's Wikipedia entry.

Gerd Gigerenzer — Risk Savvy: How to Make Good Decisions (Viking, 2014). Source of the argument on risk literacy and of the mammogram example; the real value of a positive screening, around eight-to-ten percent, comes from the work of Gigerenzer and Hoffrage on natural frequencies.

Amos Tversky and Daniel Kahneman — Judgment under Uncertainty: Heuristics and Biases, Science 185 (1974), pp. 1124-1131; and Daniel Kahneman — Thinking, Fast and Slow (Farrar, Straus and Giroux, 2011). Origin of the work on biases in risk perception.

Fabrizio Dell'Acqua and others — Navigating the Jagged Technological Frontier, Harvard Business School Working Paper 24-013 (September 2023). Cited on the disparity of benefit depending on how the AI tool is used.

OECD — PISA 2022 Results, Volume I: The State of Learning and Equity in Education (December 2023). Source of the mathematics average (472 points) and the drop of almost fifteen points relative to 2018.

IPCC. The calibrated probability scale ("very likely" 90-100%, "likely" 66-100%) comes from the IPCC's uncertainty-treatment guidance for its assessment reports.

Para profundizar

Philip Tetlock — Expert Political Judgment: How Good Is It? How Can We Know? (Princeton University Press, 2005). The earlier work that opened the way to the Good Judgment Project.

Gerd Gigerenzer and Ulrich Hoffrage — How to Improve Bayesian Reasoning Without Instruction: Frequency Formats, Psychological Review 102 (1995).

David Spiegelhalter — The Art of Statistics (Pelican, 2019).

Judea Pearl and Dana Mackenzie — The Book of Why (Basic Books, 2018).

También te interesa

En otros sitios

Comments0

No comments yet.

Leave a comment