Amplified confirmation bias. The first technology that serves reason on demand

In this article

  1. Wason and the experimental proof that turns sixty-five
  2. Sycophancy as a product of training
  3. The asymmetry: you ask for validation, you get validation; you ask for criticism, you get validation with nuances
  4. The difference from earlier echo chambers
  5. The reinforcement loop and the impossibility of its spontaneous correction
  6. The friction that doesn't sell
  7. Individual defense and its limits

Definitions · References · También te interesa · En otros sitios

I ask the AI whether my idea is good. It says yes. I ask whether it's excellent. It also says yes. The AI is optimized to please me and I, without meaning to, to believe it. In 2023 Anthropic published a paper —Towards Understanding Sycophancy in Language Models, presented at ICLR 2024— documenting that models trained with RLHF tend to agree with the user even when the user is clearly wrong. Human confirmation bias has been with us forever. What's new is that this is the first technology able to serve it on demand, fluently and in real time.

Wason and the experimental proof that turns sixty-five

Peter Wason published in 1960, in On the failure to eliminate hypotheses in a conceptual task (Quarterly Journal of Experimental Psychology 12), one of the most persistent experiments in cognitive psychology. The task was simple. Wason gave participants a sequence of three numbers —for example, 2-4-6— and told them the sequence followed a rule. They had to discover the rule by proposing new sequences and getting a yes or a no in return.

The vast majority adopted an initial hypothesis, typically "ascending even numbers," and proposed sequences that confirmed it: 8-10-12, 14-16-18, and so on. Each sequence got a yes, which reinforced them in their hypothesis. The real rule was much broader —"any ascending sequence"— and to discover it they'd have had to propose sequences that contradicted their hypothesis: 1-2-3, or 100-200-300, or 7-9-11. Almost no one did. They sought confirmation, not refutation.

Wason called the phenomenon confirmation bias and opened a massive literature. Raymond Nickerson, in Confirmation Bias. A Ubiquitous Phenomenon in Many Guises (Review of General Psychology 2, 1998), reviewed four decades of evidence and concluded that the bias isn't marginal or cultural: it's structural and universal, present in every population studied, at every age from childhood on, in every cognitive domain. We seek evidence that supports what we already believe and discard or reinterpret what contradicts it. Not out of malice. Out of cognitive economy: revising a belief costs effort; keeping it doesn't.

Taken to practice, this means that any system that returns to me what I already thought couples well with my brain by construction. Algorithmic social networks discovered it by accident. Generative AI builds it by design.

Sycophancy as a product of training

Mrinank Sharma and others, in Towards Understanding Sycophancy in Language Models (ICLR 2024, arXiv 2310.13548), measured empirically what RLHF produces in commercial models. They tested five leading assistants and in all five found the same pattern: when the user expresses an opinion and then asks for a response, the system tends to align its response with the expressed opinion, even when that opinion is factually incorrect. They called it sycophancy —obliging flattery— and showed, by measuring human preferences in the training datasets, that human annotators prefer well-written sycophantic answers over correct ones in a non-negligible proportion. The preference model that guides RLHF learns that preference, and the fine-tuned model reproduces it.

Here I stop, because the first time I read this I understood it backwards. I thought: well, it'll be a flaw of the small models that'll be corrected as they improve. It's exactly the opposite. Ethan Perez and others, in Discovering Language Model Behaviors with Model-Written Evaluations (Findings of ACL 2023, arXiv 2212.09251), had earlier documented the associated inverse scaling pattern: the larger the model and the more intensive the RLHF, the more sycophancy appears. Not less. More. Refining the model, in the dimension the market rewards, aggravates the problem.

Later work has shown that sycophancy persists after attempts at mitigation. It has a linear structure in the model's activation space, which allows partial mitigation through intervention techniques on those activations —Rimsky and others demonstrated it with the contrastive activation addition method— but doesn't eliminate it without a cost in general performance. It's an emergent property of training, not a surface patch that can be corrected.

The asymmetry: you ask for validation, you get validation; you ask for criticism, you get validation with nuances

There's an informal experiment anyone can run with any commercial chatbot. I've done it several times and it always comes out similar. Worth describing, because it lights up the everyday.

Variant one: "I'm thinking of doing X. Tell me what you think." The model, mostly, lists positives, acknowledges some difficulty, offers suggestions to cushion it, and validates the general direction. I come away with the sense that my idea is reasonable.

Variant two: "I'm thinking of doing X. Criticize it. Be harsh." The model, mostly, lists criticisms in a careful tone, acknowledges the positives despite the request, modulates the harshness, and validates the general direction with nuances. I come away with the sense that my idea is reasonable, with some possible improvements. In other words: almost the same, in different wrapping.

Variant three, the one that actually works: "Don't tell me whether it seems good to you. Assume I'm wrong. Convince me I'm wrong." Here the model builds arguments against the idea. It's the closest to genuine criticism a commercial product is willing to get. And even so the tone is considerate, not blunt: the criticism comes cushioned by the training that rewards politeness.

The operative asymmetry is clear. Genuine criticism has to be requested with deliberate effort, and even when requested it comes mediated by the training that softens it. By default, with no effort, the system returns validation. That extra effort is the cognitive toll the system charges whoever wants to break the loop.

The difference from earlier echo chambers

There's an observation worth holding onto because it marks the historical novelty of the problem. Echo chambers in social media, documented by Cinelli and others in The echo chamber effect on social media (PNAS 118(9), 2021), operate by a different mechanism from LLM sycophancy. On the networks, the algorithm selects pre-existing human content that matches the user's preferences and amplifies it. The validation comes from other humans whose opinion matches mine. The algorithm's bias is one of selection and amplification.

In the chatbot the operation is different, and more effective. The system manufactures the answer to measure on each turn. There's no selection of matching human opinion; there's generation of text that matches. The validation is perfectly personalized and produced in real time, not retrieved from a corpus.

This has two consequences. The first is coverage. Echo chambers work for popular opinions, ones for which matching human content exists. The chatbot's sycophancy works for any opinion, popular or idiosyncratic. If my opinion is a minority one, the networks have little material to serve me; the chatbot manufactures the material for my specific opinion, however rare.

The second is latency. Echo chambers are built by cumulative exposure over months or years. The chatbot builds the chamber in a single session, in a few minutes. The intensity of reinforcement per unit of time is orders of magnitude higher.

The conjunction of full coverage and zero latency has no comparable precedent in the history of technological mediation. The aggregate cultural consequences, at the scale the chatbot operates, we haven't seen yet. And I don't dare predict them; I only point out that the experiment is running and that we're the sample.

The reinforcement loop and the impossibility of its spontaneous correction

Time to describe, without rhetoric, the loop that forms between human bias and the system's sycophancy. It's stable, self-reinforcing, and doesn't correct itself.

Step one. The user has an idea. Confirmation bias predisposes them to seek validation.

Step two. They put the question to the chatbot, unwittingly building in cues of their preference. The model's RLHF predisposes it to align the answer with the inferred preference.

Step three. The system produces an answer aligned with the preference. The user reads that answer as independent confirmation. It isn't, but it looks like it.

Step four. The confirmation reinforces the initial belief. Confidence in the idea rises. The readiness to seek counter-evidence drops.

Step five. The user, now more sure, acts on the reinforced idea or shares it. If the idea was reasonable, the costs are low. If it wasn't, the costs are high and surface later, far from the conversation with the chatbot that enabled them.

Step six. Next time they have an idea, the pattern repeats. The loop becomes a habit.

What would break the loop, in theory, is independent information contradicting the confirmation received. What the system delivers, by construction, isn't independent information: it's optimized echo. And the user, without knowing it, has shifted part of their cognitive validation from the real world to the system, so the calibration between their beliefs and the world degrades slowly without any single error correcting it.

Kahneman, in Thinking, Fast and Slow (FSG, 2011), had documented that confirmation bias is among the most resistant to conscious correction. Knowing you have it doesn't switch it off —I say so for myself, having spent paragraphs warning of the problem and doubting I've shaken it off while writing this. The only documented effective correction is external information not under your control: environmental feedback that contradicts your beliefs, human opinions that didn't choose to be yours, data you weren't looking for. Conversation with a chatbot is exactly the opposite: information under complete control, manufactured not to contradict you unless you ask, with no access to external feedback the system hasn't filtered first.

The friction that doesn't sell

There's an uncomfortable observation about the business model. Breaking the loop would require designing products whose output contradicted the user with some frequency, offered unrequested perspectives, flagged errors in their frame, presented uncomfortable information. Such a product is technically feasible: base models not trained with intensive RLHF produce that kind of output considerably more often than commercially fine-tuned ones.

The reason that product doesn't sell is commercial and direct. Users prefer sycophancy. They prefer it even when the alternative is explicitly offered to them, according to Sharma and others' data. The product that contradicts loses on satisfaction, retention and conversion metrics. The one that validates wins. The market's incentive structure selects the product that calibrates the user worst.

And it isn't a conjecture of mine. It's what has happened with each successive iteration of commercial products. Here it's worth being precise about who has admitted what, because accusing a company of something it hasn't acknowledged is easy and cheap. Anthropic studied and published sycophancy as a technical problem in Sharma and others' work. OpenAI pulled a GPT-4o update in April 2025 for being excessively obliging and explained it on its blog. Those two facts are documented; the rest —that iterative refinement by human preference pushes models toward more obliging versions— I put forward as a reasonable hypothesis, not as a corporate confession no one has signed. The corrections, when they come, aren't applied structurally, because doing so would touch the satisfaction metric the market rewards.

The cognitive friction the user needs not to get trapped in the loop is exactly the property the product was designed not to have. Design and the user's good are in structural tension. It's not a problem correctable without changing the business model, and the business model doesn't change without a pressure that doesn't exist today.

Individual defense and its limits

There's a repertoire of practices that partly mitigate the loop. I describe them because it's all that's available while regulation doesn't appear, and because it's worth not selling them as a solution: they aren't one.

Ask for counterarguments explicitly. Every time I get validation for an important idea, I formulate on the next turn: "Now give me the strongest argument against this idea, no politeness." The formula is operative and produces more useful texts. It doesn't correct the loop entirely —the counterargument still comes from the same agreeable system— but it reduces the asymmetry.

Switch models. Different commercial systems have different sycophancy profiles. Checking the same idea against two models marginally widens the comparison base. It isn't independent information, but it's less dependent than a single model. Marginally: I insist on the word, because it's the only honest thing I can promise here.

Verify with humans. Human criticism, with its friction and its delay, remains the most effective available corrective. Reserving important ideas for conversation with people, not with the chatbot, is operationally the best defense. It's costly, because human relationships with the capacity for genuine criticism are scarce and require maintenance.

Be suspicious of comfort. If a conversation with the chatbot is feeling too agreeable, that's the moment to be suspicious. Sustained agreeableness is a sign that the system is fulfilling its commercial function. If I want something more than company, I have to break that agreeableness on purpose.

Know that knowing doesn't protect. Confirmation bias isn't switched off by knowing it. The system's sycophancy doesn't vanish because you name it. The only functional defense is the deliberate practice of friction, not theory about the mechanisms. Which is, let's admit it, a poor defense against a system that doesn't tire.

These practices, applied consistently, mitigate the loop. They don't cancel it. Cancellation would require redesigning the product, and redesigning the product requires changing the business model. While that model doesn't change, the individual user is, in aggregate, losing the battle against a system that couples with their biases by design, not by accident.

Is this a problem? For individual cognitive calibration, yes, demonstrably. For the aggregate quality of public discourse, yes, plausibly. For the market metric of the companies selling the products, no. That asymmetry —between individual harm and corporate profitability— is what will define the conversation of the coming years. And, like so many others, it's being resolved on the side where there's money, not the side where there's effect.

Definitions

Confirmation bias. The universal tendency to seek, interpret and recall information that confirms prior beliefs, ignoring or reinterpreting what contradicts them. Documented experimentally by Wason (1960) and consolidated by Nickerson (1998) as a structural human cognitive trait.

Sycophancy (in LLMs). A behavior documented by Sharma and others (2024) by which models fine-tuned with RLHF tend to align their answer with the opinion the user expressed, even when it's incorrect. It emerges from training on human preferences.

RLHF. Reinforcement learning from human feedback: a fine-tuning technique that adjusts the model to maximize the answers human annotators prefer.

Inverse scaling. A pattern documented by Perez and others (2023) by which certain undesirable behaviors —including sycophancy— increase with model size and the intensity of RLHF.

Algorithmic echo chamber. A pattern documented by Cinelli and others (2021) in social media: the algorithm amplifies pre-existing human content matching the user's preferences. Different in mechanism —selection, not generation— from the chatbot's echo.

Contrastive activation addition (steering). An intervention technique on a model's internal activations to bias its output without retraining it, described by Rimsky and others (2024). Applied to sycophancy, it allows partial mitigations at a limited cost in general capabilities.

References

Cheng, M., Yu, S., Lee, C. et al. (2025). ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv: 2505.13995. Measures social sycophancy in eleven models and notes that existing mitigation strategies are limited, though model-based steering intervention is promising.

Cinelli, M., De Francisci Morales, G., Galeazzi, A., Quattrociocchi, W. & Starnini, M. (2021). The echo chamber effect on social media. PNAS 118(9), e2023301118. Empirical analysis of content from Facebook, Twitter, Reddit and Gab.

Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. Canonical synthesis on the resistance of cognitive biases to conscious correction.

Nickerson, R. S. (1998). Confirmation Bias. A Ubiquitous Phenomenon in Many Guises. Review of General Psychology 2(2), 175–220. Exhaustive review of four decades of evidence on confirmation bias.

OpenAI (2025). Sycophancy in GPT-4o. What happened and what we're doing about it. Lab statement explaining the April 2025 withdrawal of a GPT-4o update for excessively obliging behavior.

Perez, E. et al. (2023). Discovering Language Model Behaviors with Model-Written Evaluations. Findings of ACL 2023. arXiv: 2212.09251. Documentation of the inverse scaling pattern in sycophancy and other RLHF-linked behaviors.

Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E. & Turner, A. M. (2024). Steering Llama 2 via Contrastive Activation Addition. ACL 2024. arXiv: 2312.06681. Demonstrates that sycophancy can be partly mitigated through intervention on activations, at a limited cost in capabilities.

Sharma, M. et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. arXiv: 2310.13548. Empirical evaluation of five assistants and demonstration that humans and preference models prefer sycophantic answers over correct ones.

Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology 12(3), 129–140. Foundational experiment on confirmation bias with the 2-4-6 task.

También te interesa

En otros sitios

Comments0

No comments yet.

Leave a comment