Invisible Biases. The Declared Bias Gets Filtered, the Baseline One Stays Intact

In this article

  1. The bias the industry does audit
  2. What no auditor puts in their report
  3. Neutral style
  4. Reliable authors
  5. Controversial topics
  6. Legal terms
  7. RLHF as the window where bias is decided
  8. The audit as a normalisation tool
  9. Style, authority, criterion. Three examples where it shows
  10. The political question, again
  11. You might also like

Definitions · References · Elsewhere

Joy Buolamwini and Timnit Gebru documented in Gender Shades (FAccT 2018) that commercial facial-recognition systems failed up to 34.7% on dark-skinned women, against barely 0.8% on light-skinned men. The industry filtered that bias, to a large extent. But a model's biases aren't the obvious ones. Those get published, audited, mitigated. The ones that hold the model up are at the base: what kind of prose is considered neutral, which authors are cited as reliable, which topics are marked as controversial. They aren't seen because they're in the decision about what to consider normal. Bias detectors detect the declared bias. The other stays intact, and the audit leaves it intact because it isn't even looking at it.

The bias the industry does audit

There's a story the industry tells with reasonable pride, and it's worth acknowledging before questioning it. The field of fairness in machine learning has spent a decade measuring differential errors by demographic subgroup. Gender Shades was a milestone because it took three commercial systems (IBM, Microsoft, Face++), evaluated them over a sample balanced by gender and by skin tone, and published the error matrix. The differences were gross and reproducible. After the paper, the three providers cut their rates within a few months. The external audit had worked.

Since then, the big labs keep internal teams measuring demographic disparities in their products. A good part of the results is published. Training is done with rebalanced data. Output filters are applied that detect racial or sexist vocabulary. The field isn't naive and the results, in the metrics it measures, are visible. Today's system is reasonably less discriminatory in classic terms than one from eight years ago.

That's the part worth accepting as is before discussing the rest. The industry isn't denying the problem. It's measuring what it knows how to measure and mitigating what it knows how to name. The serious question is what's being left unmeasured.

What no auditor puts in their report

There's a second layer of bias that operates lower down, where demographic-disparity metrics don't reach. It's made up of decisions about what's considered normal. It isn't bias in the sense of treating an identified group worse. It's a stance, assumed by default, that defines the model's stylistic, conceptual and referential centre. And the centre decides, before anything else, what counts as a reasonable question and what counts as a competent answer.

Four examples so we don't stay in the abstract.

Neutral style

The neutral style the model produces by default is that of contemporary Anglo-Saxon academic prose: medium-length sentences, controlled technical vocabulary, discourse markers like "however," "importantly," "it is worth noting," displaced into Spanish as "sin embargo," "es importante señalar," "conviene notar." When you ask the model to write "neutrally," that's the voice that comes out. It isn't neutrality, it's style. A Spanish-language writer who writes with long sentences, with subordinate clauses, with broad periods in the manner of nineteenth-century Spanish prose, will be corrected by the model as if they were verbose. Cervantes would pass an automatic style checker with a very bad mark.

Reliable authors

The reliable authors the model cites are the ones that appear frequently in the corpus. That means the big Anglo-Saxon journals and the most-cited texts in English. The model doesn't cite Spanish legal publications, nor German philosophical journals outside the three classics, nor Japanese scientific press, nor Chinese technical reports published in Beijing. It's not because it decided they're less reliable. It's because the corpus gives it less mass to distinguish them. The operational consequence is the same: the model builds, without naming it, an international canon that coincides almost perfectly with the one the Anglo-Saxon academic establishment considers its own.

Controversial topics

The controversial topics are the ones marked as such in the fine-tuning phase. The marking was done in a specific culture, with specific criteria, at a specific historical moment. A conversation about abortion provokes, in many models, a carefully equidistant answer. A conversation about the legitimacy of the British monarchy doesn't provoke it. A conversation about the rotation of work shifts doesn't provoke it. The boundary between the controversial and the neutral isn't decided by the question. It's decided by the culture of the lab that trained the model, and that culture is identifiable.

The legal terms the model knows best are those of US federal law. Ask it about a figure of Spanish commercial law, of French civil law or of Islamic fiqh, and it will compare below its average performance. Ask it about the Fourth Amendment and it will answer fluently, contextualising it even without being asked. It isn't malice, it's the composition of the corpus. But, operationally, when the model "knows about law" it knows better about US law, and the non-Anglo-Saxon user receives a bias they aren't informed of.

None of these four asymmetries appears in a fairness report of the usual kind. The standard metrics don't cover them. And that's why they can persist, intact, beneath communication campaigns that celebrate real advances in other dimensions.

RLHF as the window where bias is decided

There's a specific training phase where these decisions are taken in a concentrated way, and it's worth looking at because it's still little discussed. It's called RLHF, reinforcement learning from human feedback. After the base model has been trained on the corpus, a second phase uses human preferences to align the outputs with what's considered desirable. A group of annotators reads alternative answers to the same question and picks which they prefer. The model is adjusted to produce more often the answers the annotators preferred.

The obvious question is who annotates. Casper and others, in Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback (arXiv 2307.15217, 2023, TMLR), published an exhaustive taxonomy of the process's structural problems. They identify, among others, annotator bias as a category with no clean technical solution. Annotators are hired where qualified labour in English can be paid at a competitive price. That means, principally, Anglo-Saxon countries, the Philippines, Kenya, urban India. The annotator's culture of origin filters what they consider a "polite answer," a "clear explanation," an "appropriate tone." Those judgements enter the model as gradients.

Kirk and others, in The PRISM Alignment Dataset (NeurIPS 2024 Datasets & Benchmarks, arXiv 2404.16019), took the problem to its most complete empirical test. They recruited fifteen hundred participants from seventy-five countries and systematically measured how their preferences differed over the same answers. What they found is what's to be expected and that's why it's scandalous: preferences vary systematically by demographics. What a middle-class, university-educated US annotator considers an excellent answer isn't what a Nigerian annotator of the same age considers an excellent answer. The assistant's "default style," the tone the model produces by default when you ask it for nothing concrete, depends on who was chosen to teach it what an excellent answer is.

This is the line where the baseline bias is decided, day after day, annotation after annotation, without it appearing in any fairness report because it isn't measured as a demographic disparity of the output, but slips into the very definition of what a good output is.

The audit as a normalisation tool

Cathy O'Neil in Weapons of Math Destruction (Crown, 2016) warned of something worth rereading in light of LLMs: algorithmic tools aren't neutral, and neither are the criteria they're audited with. An audit is a look with priorities, and the audit's priorities aren't those of the universe of possible problems, they're those of the consensus of whoever audits.

The consequence is paradoxical. The more a system is audited, the more the biases the audit doesn't detect are normalised. Because the audit produces a document that says "system verified," and that document shifts public pressure toward other fronts. If I certify to you that my model doesn't discriminate by gender or race, and omit saying that my model implicitly assumes Anglo-Saxon prose as neutral, what you read is "clean model" and you stop asking.

The AI Now Institute, in its Artificial Power: 2025 Landscape Report, points out the same thing sharply. The industry has taken bias as a technical problem its internal teams can manage with internal tools. The broader problem, that of who decides what's measured and what isn't, stays outside the frame. The concentration of power in small teams of fine-tuning and annotation, all of them within a very narrow cultural band of the world, configures a baseline bias no auditor is asking to name.

Let's say it clearly. An AI "without biases," in the sense in which the industry uses the word, is an AI whose bias coincides with the auditor's. The audit doesn't eliminate bias; it homogenises it. And the homogenisation has a cost that isn't counted: the ways of speaking, of citing, of qualifying the controversial, of defining the normal, that don't fit the auditor's culture, are the ones the model will learn to avoid.

Style, authority, criterion. Three examples where it shows

There are three operational places where the baseline bias surfaces clearly enough to discuss it without going into abstraction.

First, the model that drafts "neutrally" ends up writing with translated Anglo-Saxon syntax. Short sentences, brief paragraphs, explicit transitions, conclusions that close each section. That isn't neutrality, it's the voice of US technical-writing manuals. Classic Spanish prose, academic French, philosophical German, epistolary Japanese, all of them are corrected as verbose or disorganised by a model whose criterion of "good writing" was calibrated by annotators with a single tradition behind them. It's a major stylistic bias, invisible to the user, normalised by the audit.

Second, the model that always cites the same three journals does so because the corpus has taught it that those three journals appear cited frequently. Nature, Science, The Lancet. For scientific English, they're the reference. For academic Spanish, there are dozens of relevant journals in each field the model doesn't mention because their frequency in the corpus is marginal. The authority bias works as feedback: the more the model cites the dominant journals, the more is published in them, the more they're cited in other corpora, the more the next model cites them. Canonicity self-reproduces.

Third, the model that avoids non-US legal terms doesn't avoid them by explicit order. It avoids them because when it talks about law its larger statistical mass is there. And by the logic of confidence calibration, the model prefers firm ground. The consequence is that the Spanish user who asks about prescripción adquisitiva will receive a more insecure answer than the US user who asks about adverse possession, even though they're the same legal figure with different names. And that affects real legal decision-making, now that the models are entering professional-advice workflows without the user always knowing.

The political question, again

Bender and others, in Stochastic Parrots (FAccT 2021), warned of this without yet going into industrial-scale RLHF. Crawford, in Atlas of AI (Yale UP, 2021), framed it as a geopolitical decision disguised as a technical choice. Noble, in Algorithms of Oppression (NYU Press, 2018), had documented the analogous pattern in search engines. The three readings converge on an observation worth fixing.

When bias is reduced to a technical problem, the debate stays on the surface. When it's recognised that the bias is structural, derived from who composes the corpus, who filters, who annotates, who decides what to consider normal, the debate turns uncomfortable because it has no solution from inside the lab. The solution, if it exists, runs through opening the decisions to something resembling public scrutiny, and that's the last thing the industry is willing to accept while its competitive advantages depend on secrecy.

Until that happens, the interesting figure to pin on the wall isn't that of the measurable demographic disparities, which improve. It's that of the stylistic, referential and cultural narrowing of the average model, which worsens as the industry converges on annotator training, on synthetic-data providers and on alignment best practices. A model fine-tuned by annotators from three countries speaking two languages, by metrics chosen by teams from one country, speaking one language, is a model whose neutrality coincides to the millimetre with the voice of that country and that language. Calling it objective is a rhetorical choice. What it is is consistent with who made it.

Definitions

Bias (in ML). A systematic pattern in a model's outputs that favours or disfavours a group, style, source or frame relative to others, without that preference being justified by the task. It's distinguished from simple noise by its persistence and its correlation with properties of the corpus or the alignment process.

Demographic disparity. A measurable form of bias evaluated by comparing error or behaviour rates between subgroups defined by protected variables (gender, race, age). It's what standard fairness auditing measures.

Baseline or structural bias. A less measurable form that operates over the definition of what's considered normal or neutral in the model: style, authority, relevance criterion. It usually isn't included in fairness reports because it doesn't reduce to a comparison between subgroups.

RLHF (reinforcement learning from human feedback). A training phase in which the model's answers are adjusted from annotated human preferences. Its underlying decisions (what a good answer is) are culturally situated and carry over to the model.

Annotator. A person hired to label data or answer preferences in the RLHF phase. Their cultural and linguistic profile configures the resulting model's calibration bias.

Algorithmic audit. A procedure that evaluates a system by metrics defined to detect discriminatory or harmful behaviours. Effective against the biases the metrics cover; blind to the ones they don't.

Normalisation by audit. The effect by which certifying a system shifts public pressure toward other fronts, leaving intact the biases not contemplated by the audit.

References

Buolamwini, J. & Gebru, T. (2018). Gender Shades. Intersectional Accuracy Disparities in Commercial Gender Classification. FAccT 2018. An external audit of commercial facial-recognition systems and the origin of the field of fairness applied to commercial products.

Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. DOI: . The general framework on how the corpus carries structural biases to the model.

O'Neil, C. (2016). Weapons of Math Destruction. Crown. Cited for the thesis that audits reflect the priorities of whoever does them and not necessarily the universe of problems.

Noble, S. U. (2018). Algorithms of Oppression. How Search Engines Reinforce Racism. NYU Press. Documentation of aggregate bias in algorithmic systems before the LLM era.

Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. A geopolitical framework on the composition of the corpus as a political decision.

Casper, S. et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv 2307.15217 / TMLR. An exhaustive taxonomy of the structural failures of RLHF, including annotator bias.

Kirk, H. R. et al. (2024). The PRISM Alignment Dataset. NeurIPS 2024 Datasets & Benchmarks. arXiv 2404.16019. Empirical evidence with fifteen hundred annotators from seventy-five countries that human preferences in RLHF vary systematically by demographics.

Brennan, K., Kak, A. & Myers West, S. (AI Now Institute) (2025). Artificial Power. 2025 Landscape Report (June 2025). A critique of how the industry reduces bias to a technical problem while concentrating the real decision power in small teams.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment