Statistics as a Substitute for Knowledge. The Right Answer With No One Who Knows

In this article

  1. What a model does, slowly, when it speaks
  2. Knowledge, let's recall what it was
  3. Frankfurt and bullshit, before the buzzword
  4. Three figures facing the truth
  5. What makes the difference visible
  6. Specialisation and rarity
  7. When it invents
  8. Asymmetric epistemology
  9. The operational consequence
  10. The responsibility that doesn't appear
  11. You might also like

Definitions · References · Elsewhere

An AI doesn't know anything. It calculates the most probable next word. That the answer is correct doesn't imply knowledge behind it, it implies that the calculation gave the right word in that context. Harry Frankfurt called this "bullshit" in 2005, long before the word was applied to AI: speech indifferent to truth, neither lying nor truthful. The difference from a human who knows is invisible while everything goes well. It becomes visible only when it fails, and by then you've already made the decision.

What a model does, slowly, when it speaks

It's worth being clear about the operation, because most of the arguments that follow depend on not losing it. An LLM produces text one token at a time. For each token, it calculates over the entire vocabulary a probability distribution conditioned by everything it has seen up to that point, within its context window. That distribution assigns each possible word a number between zero and one, and then a sampling mechanism, with or without temperature, picks one. It repeats the operation with the output concatenated. It repeats. It repeats. Until the system decides to cut.

There is, in that process, no internal representation of the fact that something is true. There are frequencies, there are patterns, there are neighbourhoods in an embedding space. The frequencies have been calibrated during training over trillions of tokens, and have produced a very fine distribution that, in a great many contexts, predicts as the most probable token the word an informed human would also write. But the mechanism isn't knowing. It's estimating what comes next.

The distinction isn't rhetorical. It has a very concrete consequence: when the corpus median is wrong on a point, the model is wrong along with it, and it does so with exactly the same fluency with which it gets it right when the median was right.

Knowledge, let's recall what it was

There's a classic definition we drag along from Plato. Knowledge is justified true belief. A proposition is known if it's believed, if it's true and if the belief has an adequate justification. The definition held for centuries. In 1963, in a brief three-page text, Edmund Gettier broke it with a couple of counterexamples in Is Justified True Belief Knowledge? (Analysis 23(6)). He showed that you can have justified true belief and still, intuitively, not be knowing. His counterexamples opened a debate that's still open sixty-three years later, with dozens of proposed repairs and none definitive.

What's interesting for the question about LLMs isn't the dispute among Gettier's heirs. It's that, whatever repair you accept, all the serious versions of the concept of knowledge require at least two things an LLM doesn't have. First, a subject that holds the belief, with something resembling an internal representation of the fact. Second, a justification that connects that belief with the world, not just with the corpus.

An LLM doesn't believe. What it produces isn't a belief of its own, it's a token sampled from a conditioned distribution. And justification, in the epistemological sense, is absent. The model can't explain why it asserts what it asserts, except in terms of statistical co-occurrence with the context. When asked to "justify," it produces more text that looks like justification, generated by the same mechanism: the most probable word given the context. It's synthetic justification, not anchored in facts.

Saying an LLM knows something is, therefore, an elastic way of talking. What the system has is statistical availability. The difference is invisible in a bar conversation, where statistical availability is enough to seem informed. It's decisive in any domain where trust depends on the causal chain between what's said and what happens outside.

Frankfurt and bullshit, before the buzzword

Harry Frankfurt published in 1986 an essay in Raritan, expanded into a book in 2005 under the title On Bullshit (Princeton University Press). It's a brief, dense text, technical in the philosophical sense of the word. Its central thesis distinguishes three figures.

The liar knows the truth and deliberately hides it. He has an intense relationship with the truth: he respects it enough to hide from it. The truthful person believes what they say and strives for it to be true. The bullshitter is neither. He's someone indifferent to truth. He doesn't care whether what he says is true or false. What he cares about is the effect of the speech: to convince, to keep a conversation going, to seem competent, to sell something. Truth and falsehood aren't his criteria. His criterion is rhetorical effectiveness.

Three figures facing the truth

Frankfurt closed the essay with an uncomfortable observation: bullshit is more dangerous to truth than direct lying. Because the liar at least has a map of what's true and what isn't, while the bullshitter has switched off the question. And because the culture of bullshit, when it spreads, erodes the collective capacity to tell the true from the false.

The application to LLMs is direct, and it's been seen coming for years. In 2024 a paper even appeared titled ChatGPT is Bullshit (Hicks, Humphries and Slater, in Ethics and Information Technology), arguing the equivalence with formal arguments. What an LLM does isn't lying. Nor is it telling the truth. It's generating the statistically most fitting text for the context. If that text coincides with the truth, perfect. If not, perfect all the same. The mechanism is indifferent to the question. It's bullshit in Frankfurt's technical sense, with no moral connotation, descriptively.

Putting it that way doesn't imply LLMs aren't useful. It implies their usefulness shouldn't be confused with their truthfulness.

What makes the difference visible

When everything goes well, the difference between knowledge and statistical availability is undetectable. If you ask for the capital of France, it doesn't matter how the system "knows" it. The capital is Paris, the output is Paris, the result is satisfactory.

The difference becomes visible in four typical scenarios.

The first is the specialised domain. When a question demands niche knowledge, the model has less statistical mass to train on. The answers become more fluent in form and more unstable in content. In rare medicine, in the law of small jurisdictions, in minor regional history, the LLM produces answers that sound good and that a professional in the field identifies as composited. Composited in the sense that they mix plausible elements but don't answer a specific verifiable truth.

Specialisation and rarity

The second is the unusual context. Questions where the combination of variables departs from the corpus's usual distribution. The model extrapolates by similarity, and the outputs are those someone would produce who knows the genre of the answer without knowing the specific answer. When pressed, the model doesn't correct with an "I don't know"; it corrects with another production of the same genre, equally fluent.

The third is the question where the corpus median is wrong. If online there are ten sources asserting one thing and two asserting another, and it turns out the two correct ones are the ones that arrive late to mass coverage, the model will lean toward the ten. The corpus's statistical truth can diverge from the truth. When it diverges, the model doesn't side with the second. It sides with the majority, with the same confidence as when the majority was right.

When it invents

The fourth is statistical invention (what the industry calls, without qualification, hallucination; it's worth keeping the nuance discussed in another article of this series). When the model lacks information in the corpus about a specific question and still produces an answer, it doesn't announce the deficit. It generates what would statistically look like an answer of that genre, with fabricated details that have all the texture of the real. Invented authors' names, titles of papers that don't exist, court rulings that were never handed down. The invention is bullshit in its pure state: indifference to truth kept afloat by the machinery of fluency.

In none of these four scenarios does the system warn the user. The output has the same format, the same fluency, the same confident tone as when it gets it right.

Asymmetric epistemology

There's a point worth underlining because it's often overlooked. A human who knows can be wrong, and when they are, their error usually comes with recognisable markers: hesitation, a change of tone, self-criticism if they have time to think it over. A human who doesn't know but dares to answer, too, normally, leaves traces. Human conversational culture has developed fine signals to tell informed confidence from improvisation.

An LLM doesn't have that modulation, except when it's been explicitly trained to simulate it. And when trained to seem humble, it's humble uniformly, not conditioned on the real quality of the answer. The calibration between confidence and correctness is, at best, partial: the model is just as sure when it gets it right as when it's wrong, except for specific interventions the industry hasn't yet generalised.

The consequence is asymmetric. The user evaluates the model with the criteria they'd use for a human. Fluency is read as command of the topic. The absence of hesitation is read as well-founded certainty. The length of the answer is read as depth. Those three markers, which in a human correlate reasonably with the quality of knowing, in an LLM correlate with nothing. They're traits of the model's style, not of the content.

The operational consequence

When someone delegates to an LLM a task whose correctness matters, they're trusting the corpus's statistics, not knowledge. This substitution is invisible in the interface, is masked by fluency, and appears in no usage instruction of the product. The user doesn't see it and has no way of seeing it except when something fails spectacularly and forces them to open the hood.

The responsibility that doesn't appear

The distinction matters because it underpins different decisions. Trusting the knowledge of a human expert admits responsibility: if they're wrong, there's someone who answers for it, there's a correction mechanism, there's reputation at stake. Trusting the corpus's statistics admits no such responsibility. The corpus isn't wrong, it simply reflects what was there. The model isn't wrong, it simply predicts what comes next. And the provider company, in its terms of service, expressly declines any guarantee over the correctness of the outputs. They say it themselves in writing, in print the user rarely reads.

Bender and others, in On the Dangers of Stochastic Parrots (FAccT 2021), described LLMs as systems that generate plausible sequences of tokens with no access to meaning. The stochastic-parrot metaphor has aged well. What an LLM does when it speaks is statistically sophisticated, computationally impressive, and epistemologically indistinguishable from Frankfurt's bullshit.

Bender and Koller, in Climbing towards NLU (ACL 2020), had already laid down the conceptual argument: a system that learns about form with no access to meaning can't aspire to understand, only to imitate. Marcus and Davis, in Rebooting AI (2019), earlier picked up the criticism of using epistemic vocabulary for systems with no explicit representation of facts. Hahn and Goyal in A Theory of Emergent In-Context Learning as Implicit Structure Induction (arXiv 2303.07971, 2023) offer a somewhat kinder reading. They suggest that, over sufficiently compositional data, the model ends up inducing latent structures that recombine operations. Xie and others, in An Explanation of In-context Learning as Implicit Bayesian Inference (ICLR 2022), reach a parallel conclusion: ICL can be seen as a Bayesian inference over latent tasks implicit in the corpus. The two readings raise the ceiling of what statistics can simulate. Neither of them dissolves the difference between simulating the result of a reasoning and having reasoned it.

And when the simulation is enough, it's worth having a proper name for what's happening. The name, since Frankfurt, already exists.

Definitions

Token. The minimal unit of text a model processes. Roughly half a word in English. LLMs operate on sequences of tokens, not on words or letters.

Next-token prediction. The central operation of an LLM. Given an input sequence, the model calculates over the entire vocabulary a probability distribution for the next token, and samples one. The operation repeats until generation is cut.

Knowledge (classical definition). Justified true belief. A Platonic definition refined up to the twentieth century, when Gettier broke it with counterexamples. Any serious repair still requires a subject, belief and justification, none present in the strict sense in an LLM.

Gettier case. A situation in which someone has a true and justified belief and yet, intuitively, doesn't know. It showed the classical definition of knowledge was insufficient and opened contemporary epistemology.

Bullshit. A technical concept introduced by Harry Frankfurt in 1986 and formalised in 2005. Speech indifferent to truth, distinguishable both from lying (which hides it) and from truthfulness (which pursues it). It's defined by the speaker's attitude toward the question of whether what's said is true.

Stochastic parrot. A metaphor popularised by Bender, Gebru, McMillan-Major and Shmitchell (2021) for LLMs. Systems that generate plausible sequences of tokens with no access to the meaning of what they produce.

Hallucination. The production by an LLM of information that looks factual but is false: invented names, nonexistent citations, fabricated data. It's a particularly visible expression of the mechanism's bullshit character.

Calibration. The correspondence between the confidence with which a system asserts something and the real probability that it's true. In LLMs, calibration tends to be poor: fluency doesn't vary with the degree of factual backing.

References

Frankfurt, H. G. (2005). On Bullshit. Princeton University Press. Original essay in Raritan (1986), expanded into a book. The operational distinction between lying and bullshit as distinct epistemic categories.

Gettier, E. L. (1963). Is Justified True Belief Knowledge?. Analysis 23(6), 121–123. Classic counterexamples to the Platonic definition of knowledge, cited as proof that the concept requires more than an LLM can offer.

Hicks, M. T., Humphries, J. & Slater, J. (2024). ChatGPT is bullshit. Ethics and Information Technology 26, 38. A formal application of Frankfurt's concept to LLMs, cited for its argumentative clarity.

Bender, E. M. & Koller, A. (2020). Climbing towards NLU. On Meaning, Form, and Understanding in the Age of Data. ACL 2020. The distinction between form and meaning in language systems.

Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. DOI: . The origin of the stochastic-parrot metaphor.

Hahn, M. & Goyal, N. (2023). A Theory of Emergent In-Context Learning as Implicit Structure Induction. arXiv 2303.07971. A theoretical bound linking next-token prediction with structure induction over compositional data.

Xie, S. M. et al. (2022). An Explanation of In-context Learning as Implicit Bayesian Inference. ICLR 2022. arXiv 2111.02080. A reading of ICL as Bayesian inference over latent tasks in the corpus.

Marcus, G. & Davis, E. (2019). Rebooting AI. Building Artificial Intelligence We Can Trust. Pantheon. An early critique of using epistemic vocabulary for systems with no explicit representation of facts.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment