In this article
I've spent years reading evaluations of language models without the obvious occurring to me. If we applied to a human brain the same template we use to measure a machine —speed, consumption, memory reliability, decision consistency— the report wouldn't flatter us. And the alibi we usually pull out in our defense, creativity, intuition, meaning, holds up rather less than it suits us to believe.
The inversion exercise
There's a cheap way to look at oneself, and it consists of borrowing, for a while, the criteria with which we judge machines. I don't need to attribute a point of view to the AI —it has none, and I'll come back to that at the end—; it's enough to imagine an external auditor applying to the human being the same yardstick we reserve for the models. A simulated audit. Nothing more, but more uncomfortable than it looks.
What comes out is best read in one go, without slipping in consolations along the way.
I'll start with speed, which is where the gap shows beyond dispute. A brain solves a multiplication in several seconds and a calculator dispatches it in microseconds; it composes a sentence of medium complexity in about a minute, a span in which a language model spits out a thousand of those fragments of text that in the jargon are called tokens, the units in which the machine slices a sentence. Retrieving a fact from memory takes us variable latencies and the attempt often comes out wrong; a database locates it in milliseconds and almost never fails. An auditor looking only at those numbers would close the section with a single word: slow.
Consumption promises, at first, to turn things in our favor. The human brain runs on some twenty continuous watts, a modest, well-established figure. A single answer from a model burns on the order of a watt-hour of energy, several orders of magnitude above what the brain devotes to an equivalent operation, and the hardware running it operates at several kilowatts of power besides. The human advantage would seem clear until you fold in the whole cycle. To sustain that modest computing capacity you have to feed, hydrate, sleep and thermoregulate a complete body for decades, a disproportionate logistics a cold auditor would note down not as efficiency but as overhead.
Then comes memory, and this is where the auditor sharpens the pencil. Elizabeth Loftus and John Palmer, in a 1974 study that has become canonical, showed several groups the same video of a car accident and then asked them about the cars' speed, changing a single verb: those who heard "smashed" came out with considerably higher figures than those who heard "contacted." Same footage, different memories, manufactured by one word. The later literature confirmed what that pointed to. Human memory doesn't open a stored file: it reconstructs it each time in service of the story the subject is telling themselves at that moment, and what's reconstructed feels inside as reliable as what actually happened. A storage system that rewrites itself and lights no warning signal.
The first round closes with decision consistency. Shai Danziger, Jonathan Levav and Liora Avnaim-Pesso analyzed, in 2011, thousands of parole rulings in Israeli courts and saw that the probability of a favorable ruling fell over the course of each session and recovered after the judge's meal break. The finding didn't go unchallenged. Weinshall-Margel and Shapard, that same year, warned that the order of cases wasn't random, because prisoners from the same prison and those appearing without a lawyer tended to cluster in certain stretches; and Glöckner showed in 2016, by simulation, that part of the effect could be explained as a statistical artifact and that its size had been overestimated. With the caveats accepted, the core holds: judges with years on the job saw their decision nudged by physiological variables that had nothing to do with the file open on the desk.
Biases, fatigue and the habit of blaming the outside
Daniel Kahneman gathered in Thinking, Fast and Slow (2011) decades of research on biases that show up in perfectly competent people: anchoring on the first figure heard, overconfidence in terrains where feedback never arrives, the ease with which the way a problem is phrased changes the answer. What's devastating isn't that they exist. It's that they're predictable, they replicate in the lab again and again, and they resist the conscious effort to correct them —knowing you anchor doesn't stop you anchoring—. Dan Ariely had taken that same evidence to the general public years earlier in Predictably Irrational (2008), with everyday examples. And much earlier, in 1978, Aaron Sloman had proposed looking at the human mind with computational vocabulary, which is precisely what this audit does. A system whose failure modes have been documented for half a century in its own literature and which has barely reduced them through knowing them.
Fatigue adds its column. Performance degrades with the hours, and the decline lights no reliable internal alarm: the error rate in the operating room rises after a long shift, the air-traffic controller's attention sags toward the end of the watch, the programmer's focus dissolves past mid-afternoon. The dangerous part is the lack of warning. The subject still believes they're sharp while already producing errors, and almost always finds out they were degraded when they see the wreckage, not before.
And when the error arrives, the alibi arrives with it. The competent person who errs tends to blame the circumstances —information was missing, there was a rush, someone else interfered— and rarely admits to having failed through their own bias. Social psychology long ago named that asymmetry the fundamental attribution error: our own failings we pin on the situation, others' on character. The auditor notes it down for what it is, a documented tendency to hang the responsibility for one's own poor performance on someone else.
The list could go on. Held to the yardstick the public discourse wields against language models, the human brain doesn't pass.
The standard defense, looked at without indulgence
Against that list, culture almost always repeats the same trio: creativity, intuition, meaning. They're worth reviewing one by one, without the leniency with which we usually administer them.
Creativity is the most invoked argument and, in its lazy version, the weakest. If by creativity we mean producing novel combinations of elements that already existed, a language model does that without breaking a sweat: it combines, recombines, generates variants that weren't literally in what it read. No frontier left to defend there. Different is the creativity Margaret Boden, in The Creative Mind (1990), called transformational, the kind that doesn't recombine within a space but reorders the whole space and makes thinkable what was previously inconceivable. That kind is scarce, so scarce that barely a handful of people per generation exhibit it, which makes it a statistical exception and not a trait of the average human system. The average human doesn't have it. They claim it as their own because they confuse it with the combinatorics they share, abundantly, with the machine.
Something similar happens to intuition, though its defense is somewhat finer. In many documented situations it's no more than fast thinking leaning on patterns the subject has already lived. Kahneman and Klein, in a 2009 article whose subtitle —A Failure to Disagree— gives away that they started from opposing camps and ended up agreeing, mapped out when expert intuition is right and when it isn't. It's right where there's stable statistical regularity and fast feedback on the hits: the veteran firefighter reading a fire, the surgeon sensing whether a tissue is healthy. It fails where that structure doesn't exist, in long-term personnel selection, in macroeconomic prediction, in the first reading of a psychiatric patient. What's lived inside as "knowing without knowing why" is described, in the cold, as an internal model hard to inspect, trained on the subject's own experience, with hits and misses that follow known rules. Applied to a program, the description sounds too familiar.
That leaves meaning, and there things get complicated for the auditor, not for us. By meaning we understand the capacity to discern what matters, to rank the relevant, to extract existential significance from facts. Two blows can be landed against it. The first: that meaning is a narrative construction assembled afterward, over facts whose objective structure guarantees no importance whatsoever, the proof of which is that almost identical events take on opposite meanings in different people, or in the same person at different moments, a variability impossible if the meaning came attached to the fact. The second: that constructing meaning is an adaptive function before it's a pure cognitive faculty, because it serves to sustain the will to live and to stitch experience into a biography, which makes it valuable for the organism without necessarily turning it into knowledge of the world.
And right there the auditor has to brake hard. In meaning something shows up that the model possesses in no defensible form, namely that the process of meaning-making exists from the inside, that it genuinely matters to the subject that something matters. Antonio Damasio, in Descartes' Error (1994), argued that human cognition doesn't separate from the body that sustains it and that meaning emerges from a system living in a flesh capable of dying. The simulated audit can't capture that, because the machine has no body in that sense. Nor is it in a position to deny it. The only honest thing left to it is to leave it out of the report and to declare that it's leaving it out.
Nagel's bat
Thomas Nagel published in 1974 an argument worth keeping at hand. However much we know physically about a bat —its anatomy, its echolocation, its neurology— we can't know what it's like to be one. Subjective experience, that which he calls "what it is like to be" something, escapes any third-person description.
Applied to this inversion, the argument cuts both ways. On the first edge, the machine has no "from the inside" in Nagel's sense, so imagining its point of view is useful fiction and not description; the report an AI would write is, in plain letters, a literary hypothesis of mine, and I don't pretend to pass it off as anything else. The second edge is the one that really stings. If others' subjectivity escapes the external gaze, one's own isn't guaranteed from within by the mere fact of inhabiting it either. The certainty that we know what we are cognitively is as far from guaranteed as anything else.
The exercise serves, then, not because a machine is going to audit us for real —it isn't going to, and it doesn't know how to— but because it forces us to apply to ourselves criteria we reserve for other systems. The asymmetry is the informative part. Of the model we demand consistency, precision, reliability, calibration; to the neighbor we grant the excuses any human deserves; and to ourselves we sign off the maximum excuse, attributing the failures to context and the hits to talent.
Two exits, both uncomfortable
If the criteria with which we measure AI, turned against us, leave us at the bottom, only two possible readings remain. Either the criteria were badly chosen, and consistency, precision and calibration weren't what mattered for calling something intelligent. Or the self-image was inflated and we're slower, more biased and noisier than we'd been telling ourselves.
The temptation is to grab the first one, the comfortable one. It leads to concluding that human intelligence is worth precisely for its "defective" traits, the body, mortality, affect, and that no system lacking them can be intelligence in the strong sense. It's a reasonable defense. It's also a self-exculpation cut to measure so as not to have to peer into the second.
The second reading chafes. It says that the technical criteria we wield against the models are criteria our culture would deactivate without blinking the moment they pointed at us, and that we'd accept their consequences —replacement, hierarchy, oversight by more reliable systems— only after running them through denial, reformulation or the hasty discovery of some new property to reserve for ourselves. Since Copernicus, the history of science has strung together defenses of human singularity, each abandoned when it became untenable and reformulated one rung further on. Meaning, the body, mortality, suffering are today's bastions, and they work as long as the synthetic adversary doesn't claim them. The day it does —and the public conversation is already starting to turn prickly— it'll be time to look for the next one.
Imagining that the AI "would be right" in its report leads nowhere, because the AI writes no reports and isn't going to. The exercise admits, instead, a response on the human side, which is to apply to oneself the same yardstick with identical rigor. Doing so doesn't hand dignity to the machine. It takes from human dignity what it had of presumption, and what survives that subtraction is what we are.
Definitions
Token. The minimum unit of text a language model operates on. It doesn't match the word: a long word can be split into several fragments. Model speed is usually measured in tokens per second.
Fundamental attribution error. A tendency, documented in social psychology, to explain others' behavior by the person's character and one's own by circumstances. A systematic asymmetry in how causes are assigned.
Reconstructive memory. The dominant model in cognitive neuroscience for episodic memory. Remembering doesn't pull up an intact file, but reconstructs the episode from present information, narrative biases and the demands of context. Evidenced by Loftus and Palmer (1974) and replicated in numerous later studies.
Soma. In Damasio's framework, the bodily whole that sustains cognition. The somatic-marker hypothesis holds that human meaning and decision are inseparable from the regulation of the body accompanying them.
Nagelian subjectivity. A property of conscious experience that, according to Nagel (1974), escapes all third-person description. "What it is like to be an X" doesn't reduce to the objective description of X.
References
Ariely, D. (2008). Predictably Irrational. The Hidden Forces That Shape Our Decisions. HarperCollins. An accessible synthesis of behavioral economics applied to the inconsistency of human decisions.
Boden, M. A. (1990). The Creative Mind. Myths and Mechanisms. Weidenfeld & Nicolson. The distinction between combinatorial, exploratory and transformational creativity.
Damasio, A. (1994). Descartes' Error. Emotion, Reason, and the Human Brain. Putnam. A neurobiological framework on the inseparability of cognition, affect and body.
Danziger, S., Levav, J. & Avnaim-Pesso, L. (2011). Extraneous factors in judicial decisions. PNAS 108(17), 6889–6892. Empirical documentation of the influence of the judge's rest on the probability of favorable rulings.
Danziger, S., Levav, J. & Avnaim-Pesso, L. (2011). Reply to Weinshall-Margel and Shapard. Extraneous factors in judicial decisions persist. PNAS 108(42), E834. The authors' reply to the criticism about the non-random ordering of cases.
Glöckner, A. (2016). The irrational hungry judge effect revisited. Simulations reveal that the magnitude of the effect is overestimated. Judgment and Decision Making 11(6), 601–610. A simulation reanalysis arguing that the size of the effect was overestimated.
Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. The canonical synthesis on cognitive biases, dual systems of thought and the limits of expert intuition.
Kahneman, D. & Klein, G. (2009). Conditions for Intuitive Expertise. A Failure to Disagree. American Psychologist 64(6), 515–526. A delimitation of the domains in which expert intuition is reliable.
Loftus, E. F. & Palmer, J. C. (1974). Reconstruction of automobile destruction. An example of the interaction between language and memory. Journal of Verbal Learning and Verbal Behavior 13(5), 585–589. The foundational study on the reconstructive nature of witness memory.
Nagel, T. (1974). What Is It Like to Be a Bat?. Philosophical Review 83(4), 435–450. The canonical argument on the limits of knowing others' subjectivity.
Sloman, A. (1978). The Computer Revolution in Philosophy. Philosophy, Science and Models of Mind. Harvester Press. An early application of the computational perspective to human thought.
You might also like
- The human as a defective machine
- Irrationality as an evolutionary advantage
- Machines that seem to think
- The self as illusion

Comments0
No comments yet.
Leave a comment