The Illusion of Neutrality. A Neutral AI Is an AI Whose Bias Matches Yours

In this article

  1. The chain the word "neutral" tries to hide
  2. Why language can't be neutral
  3. Perceptual asymmetry, or why no one sees their own bias
  4. All the axes, not just one
  5. When the bias changes, it showed
  6. Rozado: the bias appears with RLHF
  7. Perez: seeming opinionated sells
  8. The distance between opacity and secrecy
  9. The critique is of the label, not the intention
  10. You might also like

Definitions · References · Elsewhere

"This AI is neutral, it only states the facts." False by construction. The relevant facts are selected by someone, the tone frames them by someone, the corpus weights them by someone, the RLHF polishes them by someone. A neutral AI is an AI whose bias matches yours. When you find one that doesn't, you call it biased. Neutrality as a product attribute is the oldest cliché in tech marketing, and the most effective, because it looks descriptive instead of aspirational.

The chain the word "neutral" tries to hide

Before a model answers a user's first question, at least six human decisions have occurred that configured its voice. It's worth listing them not as an academic exercise but so the idea of "neutrality" lands where it has to be: in logical impossibility.

First decision: which corpus was used for training. Linguistic composition, temporal distribution, sources included, sources absent. A specific team decides that under legal, commercial and technical constraints. Second: which filters were applied to the corpus. What was removed, with which blocklists, with which quality criteria, with which operational definition of toxicity. Third: which architecture was chosen and with which hyperparameters it was trained, which influences what kind of patterns the model finds most easily. Fourth: which fine-tuning was done after the base training, with which specific corpus, with which weight. Fifth: which RLHF was applied, with which annotators, with which definition of a preferable answer. Sixth: which system prompt accompanies each inference, which post-hoc moderation filters certain outputs, which hard rules veto certain topics.

Calling the final product of that chain "neutral" is like calling neutral a press editorial signed by a specific author under a specific editorial line. The chain of decisions can be reasonable, can be careful, can aspire to impartiality. It can't be neutral. Neutrality, in systems that generate language, isn't an attainable attribute. It's marketing.

Why language can't be neutral

There's a deeper argument, prior to the technical one, worth putting in its place. Language doesn't transmit information about the world from a point of view of nowhere. It encodes perspective. When you say "migrant workers" or "illegal immigrants," the information about the person referred to is the same, and the implicit valuation is opposite. When you say "regime" or "government," the factual content doesn't change, and the moral frame does. When you say "liberator" or "warlord," the historical subject may be the same, and the reading falls to one or the other side of the boundary.

This isn't a discovery of contemporary critical analysis. It's what any attentive reader of Aristotle, of Machiavelli, of Orwell or of any classical rhetoric manual finds in the first lesson. Language persuades even when its author doesn't mean to persuade. The selection of terms, the order of clauses, the implicit hierarchy of what's mentioned first and what's taken as obvious, are unavoidable operations, and all of them come marked by the discursive tradition the speaker comes from.

An LLM that generates language doesn't escape this condition. It inherits it. Each answer has a lexical, syntactic and narrative texture that is the weighted median of the corpus that trained it and of the fine-tuning that polished it. That median is a stance, conscious or not. Asking the model for a neutral answer asks it to produce what its training presented to it as neutral style, and that neutral style is the voice of a specific group at a specific moment. There's no neutrality, there's a cultural convention about what counts as such.

Perceptual asymmetry, or why no one sees their own bias

Here comes an uncomfortable cognitive property that affects both users and auditors. We detect bias when it differs from ours. We're blind when it coincides.

It's the modern version of the old proverb about the speck and the plank. When a model answers a political question with a stance that departs from what the reader considers obvious, the reader marks it as biased. When it answers with a stance that coincides with the reader's, they mark it as objective. It isn't a dishonest operation. It's the default operation of the human cognitive apparatus, and it happens to all of us.

The operational consequence is that public discourse about LLM bias moves in situational lurches. Every time a specific model produces an answer that clashes with the sensibility of a powerful group, that group denounces the bias, the company adjusts, the model changes. The groups whose sensibility did coincide with the model's output noticed nothing before and notice the change afterward as if the model "had started being biased," when what's happening is that the bias has shifted in their direction of clashing.

All the axes, not just one

This holds for all the axes. It holds for the progressive sensibility that complained about Stable Diffusion generating white doctors by default. It holds for the conservative sensibility that complained about Gemini generating black vikings in historical images. It holds for the nationalist sensibility that complains a model doesn't respect its version of the territorial conflict. In all three cases, what's happening is the same: the model has been calibrated in one direction, and the groups whose bias coincides with that direction don't perceive it while the groups whose bias differs denounce it.

The reasonable conclusion isn't to pick a side. It's to understand that the word "neutral" in this context doesn't name a property of the model. It names a coincidence between the model and the observer.

When the bias changes, it showed

There are three recent cases worth bringing in because they make the chain explicit, against the industry's wish to keep it in the shadows.

The first, the Gemini incident with the generation of historical images in February 2024. Google's model produced images of racially diverse Nazi soldiers, black vikings and female popes in response to requests for historical figures. The company, faced with the scandal, suspended the generation of people, acknowledged the problem and apologised. The important thing about the case isn't the scandal itself. It's the implicit admission: the model had been fine-tuned to introduce demographic diversity by default into images, and that decision, taken by a specific team with identifiable intentions, had produced a bias opposite to historical reality. The "neutrality" the model presented before was a deliberate calibration. When the calibration changed, the bias became visible. While the calibration was invisible, there was no scandal.

Rozado: the bias appears with RLHF

The second, the quantitative work of David Rozado. In The Political Bias of ChatGPT (Social Sciences, 2023) and then in The Political Preferences of LLMs (PLOS ONE 19(7), 2024, arXiv 2402.01789), Rozado applied standardised political-orientation tests to a growing set of models. His conclusion, replicated by third parties, was that most commercial LLMs answer with a left-of-centre bias on standard political questions. What's interesting methodologically is that the base models, before RLHF, didn't show that bias so markedly. The bias, according to the study's data, sharpens during the alignment phase. It isn't an artefact of the corpus alone: it's an effect of the human process of polishing the outputs.

Perez: seeming opinionated sells

The third, the work of Perez and others, Discovering Language Model Behaviors with Model-Written Evaluations (Findings of ACL 2023, arXiv 2212.09251). They showed that more RLHF tends to produce stronger declared stances on concrete political questions, like the debate over guns or immigration. The authors' interpretive hypothesis is interesting: RLHF pushes the model to seem opinionated because opinionated answers are perceived as more useful by the annotators. The pursuit of usefulness ends up, without anyone having asked for anything like it, politically positioning the model in the direction of the median of its annotators.

The three cases share a structure. There's a human decision behind it. The decision has an identifiable direction. When that direction coincides with the observer's, the model is called neutral. When it differs, it's called biased. The word "neutral" acts as paint on top.

The distance between opacity and secrecy

There's a nuance worth introducing before closing. Not all opacity is malicious secrecy. A good part of the teams that fine-tune commercial models work with criteria they themselves consider honest: reduce visible harm, not support extreme discourse, respect the legality of the jurisdiction they operate in. Saying their model isn't neutral doesn't imply they're cheats. It implies that the result of their decisions, summed, configures a voice that's nobody's voice, but that shares identifiable codes with the place and culture it operates from.

The critique is of the label, not the intention

The serious critique isn't of intentionality. It's of the label. If instead of selling the product as "neutral" it were sold as "a model fine-tuned by a Silicon Valley team with contemporary Western moderation criteria," the cognitive contract with the user would be cleaner. The user would know what they're reading. They could compare it with another model fine-tuned somewhere else with other criteria. They could use several models for different purposes. The word "neutral," as the only label, cancels that comparison because it sets the product as the measuring stick for the rest, and everything that differs from it will appear as a deviation.

That's why the question that matters isn't "is this model neutral?", a question to which the only honest answer is "no, no model is." The question that matters is "whose bias is the one presented to me as neutrality?". That second question does have operational answers. It has company names, team names, jurisdictions, cultural frames. It has a geography. And as long as that geography goes unnamed, the word "neutral" will keep doing the rhetorical work it's designed for: rendering invisible whoever signs the weights.

Crawford, in Atlas of AI (Yale UP, 2021), posed the question with a couple of whole chapters on the political materiality of AI. Bender and others, in Stochastic Parrots (FAccT 2021), explained why neutrality in language systems is a ghost. O'Neil, in Weapons of Math Destruction (Crown, 2016), had warned of the problem when we still talked about small predictive models. The three readings converge on the same thing. It isn't about eliminating bias, which is logically impossible. It's about declaring it, and declaring it with enough precision that the user can evaluate whether the declared bias serves them or not for what they're going to use the model for.

As long as the declaration doesn't exist, the word neutral will keep being one of the most expensive words in contemporary tech marketing. You load it as an adjective and you pay with the ceding of your judgement.

Definitions

Neutrality (in a generative system). An attribute assigned rhetorically to a model whose outputs coincide with the observer's cultural expectations. It isn't a technically attainable property: any model that generates language encodes perspective by construction.

Chain of decisions. The sequence of human choices that configure a model: corpus composition, filters, architecture, fine-tuning, RLHF, system prompt, moderation. Each link has an identifiable signature.

RLHF (reinforcement learning from human feedback). A training phase in which outputs are adjusted according to the preferences of human annotators. It's one of the main determinants of the model's final bias.

System prompt. A persistent instruction that precedes the user's queries in each inference. It configures the model's tone, restrictions and answer criteria. It's usually secret.

Perceptual asymmetry. A cognitive property by which we detect bias in others when it differs from ours and are blind when it coincides. It affects public judgement of LLMs.

Political calibration. The aggregate position of a model on standardised political-orientation tests. Empirically documented as sensitive to the RLHF process.

References

Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. A framework on the political materiality of AI systems and the impossibility of neutrality by construction.

Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. DOI: . An argument on why LLMs can't be neutral in their production of language.

O'Neil, C. (2016). Weapons of Math Destruction. Crown. An early critique of the language of objectivity applied to algorithmic systems.

Rozado, D. (2023). The Political Bias of ChatGPT. Social Sciences 12(3), 148. The first systematic quantitative study of measurable political bias in commercial LLMs.

Rozado, D. (2024). The Political Preferences of LLMs. PLOS ONE 19(7), e0306621. arXiv 2402.01789. An extension to twenty-four models with eleven tests; the left-of-centre bias is observed consistently and sharpens with RLHF relative to the base models.

Perez, E. et al. (2023). Discovering Language Model Behaviors with Model-Written Evaluations. Findings of ACL 2023. arXiv 2212.09251. Evidence that more RLHF sharpens the model's declared political stances.

The Verge, Reuters et al. (February 2024). Coverage of the Gemini incident with the generation of historical images. Reports documenting the deliberate diversity adjustment and the subsequent suspension of the feature.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment