In this article
Common Crawl, the base corpus of most open models, runs at around forty percent English-language content according to the organisation's own public statistics —the exact figure dances from one monthly archive to the next— and leaves low-resource languages at minuscule percentages. What goes into training defines what the model "thinks." If the corpus is massively Anglo-Saxon, the model thinks in Anglo-Saxon and translates the rest. If you exclude certain sources, the model doesn't cite them. The choice of dataset is the quietest political choice in the sector. Those who make it don't call themselves politicians, but they decide the air the world's intelligence breathes.
The piece public discourse ignores
When the regulation of AI is debated publicly, the discussion tends to centre on two points: the model's architecture and its outputs. How much the network can do, what it says and what it shouldn't say, what safety risks it generates. There's a third piece, an intermediate one, that gets far less talk, and that decides a good part of the result long before the architecture processes anything. That piece is the training dataset.
The reason it gets little talk isn't academic. It's legal and commercial. The exact composition of the corpora of the big commercial models isn't public. What we know of the training of GPT-4, Claude or Gemini are generic mentions in marketing documents and reasonable deductions from observed behaviours. The companies don't publish the breakdown because a good part of the corpus was obtained with rights over which there are open lawsuits, and because the exact composition is a competitive advantage. The opacity is functional to the business.
Where the composition is known is in the open corpora, and that lets us read by extension what probably happens in the closed ones. And what reads back is uncomfortable if you aren't American, white, anglophone and born between 1990 and 2015.
What is documented
There are three auditing exercises worth keeping in mind so as not to talk in the abstract.
The first is Dodge, Sap, Marasović and others, in Documenting Large Webtext Corpora. A Case Study on the Colossal Clean Crawled Corpus (EMNLP 2021, arXiv 2104.08758). It documents the composition of C4, a filtered, cleaned version of Common Crawl that has trained a good part of the T5 family and its derivatives. They find three hard things. One, that blocklist filtering disproportionately excludes text about and from sexual, racial and religious minorities: the filter was designed against "toxic content" and ends up also crushing innocuous text about communities the filters inherited as suspect. Two, that an important part of the corpus contains machine-generated text, already in 2021. Three, that within the corpus appear NLP evaluation benchmarks, which means any model trained on it has seen the exams it will later be graded with.
The second is Gao, Biderman and others, in The Pile. An 800GB Dataset of Diverse Text for Language Modeling (arXiv 2101.00027, 2020). They documented in detail, for the first time at this level, what portion of the corpus came from which source: Wikipedia, GitHub, books, ArXiv, forums, the press. The Pile remains one of the few cases where you can look inside the dataset without signing an NDA, and that's why it remains an obligatory reference for understanding what goes into the statistical mass the model will use as a map of the world.
The third is Longpre and others, in The Data Provenance Initiative. A Large Scale Audit of Dataset Licensing and Attribution in AI (Nature Machine Intelligence, 2024, arXiv 2310.16787). They audited more than eighteen hundred text datasets used in practice to train and evaluate models. On the popular download repositories they found licence-omission rates above seventy percent and attribution or categorisation errors above fifty percent. It's not a woke critique nor an attack on the sector. It's an audit with figures. And the general picture is that the discipline of documenting provenance is exceptional, not standard.
Common Crawl, talking in figures
So as not to stay in abstractions, it's worth looking at what the Common Crawl Foundation itself publishes in its statistics. The figures vary by monthly archive because the corpus grows and the composition rebalances, so it's worth taking them as an order of magnitude and not as a fixed datum. In a recent archive English is the main language in around 41% of documents. Behind it, at a great distance, Russian, German, Japanese, French, Spanish and Chinese share a second band of around 5 or 6% each, in figures close to one another that dance from one archive to the next. And below, a long descent down to the dozens of languages that appear at minimal percentages and the hundreds of living languages that barely peek through, or that simply don't appear.
Read carefully. The distribution of native speakers in the world doesn't resemble that distribution of tokens. Mandarin Chinese has more native speakers than English. Hindi has more than Spanish. And in the corpus, the hierarchy is set by something else: how much written digital content is available in each language, which in turn depends on how much content has been digitised, which technological platforms have been used, which cultural hegemony has pushed publishing in one language over another. The corpus reflects the digital history of the last quarter-century, not the linguistic population of the planet.
A model trained on that distribution inherits that history. When it answers in Spanish, it does so with a Spanish it knows less deeply than English. When it opines on a regional topic outside the Anglo-Saxon axis, it does so with less statistical mass and therefore with a greater margin of improvisation. When it picks examples, it picks them from the country whose content weighs most. It's not the model's prejudice. It's the geometry of the corpus.
What Crawford called cartography
Kate Crawford, in Atlas of AI (Yale University Press, 2021), made an observation that took a while to sink in and that remains valid. The dataset, she says, isn't a neutral piece on which a neutral architecture is applied. It's cartography. And like all cartography, it reflects decisions: what's included, what scale is used, what's considered centre and what periphery. A map with the Pacific Ocean at the centre isn't the same object as a map with the Atlantic at the centre, even though both represent exactly the same sphere. The centre decides what's seen at a glance and what requires the eye to move.
The choice of dataset operates with the same logic. It decides which texts are seen at a glance and which remain in the periphery the model reaches only with effort and with growing error. It decides which authors are the corpus's default voice and which appear as the exception. It decides which countries give the examples for everything else. It decides it in silence, before the model has been trained, in a series of spreadsheets no regulator has ever publicly asked for.
The difference from an old map is that the old map was avowedly intentional. Mercator knew his projection distorted the high latitudes; he accepted it in exchange for navigational usefulness. Whoever builds an LLM's corpus doesn't declare the distortion. The distortion reaches the user as a property of the model, not as a cartographic decision.
The filtering that's also opinion
There's a second layer, subtler, where the political decision concentrates and again becomes invisible. It's the filtering of the corpus. It's not enough to choose which sources to put in; you have to decide what to take out. Raw corpora are too noisy, contain spam, contain sexually explicit content, contain violent language, contain bot-generated text. They're filtered. The decision of how to filter has almost as much importance as the decision of what to include.
What Dodge and others showed about C4 is that blocklist filtering, set up with good intentions, ends up cutting valuable text. If your blocklist contains words from the LGBTQ lexicon, you don't only remove pornography: you remove academic, journalistic and literary text that uses them descriptively. If your blocklist contains "loaded" political terms, you don't only remove propaganda: you remove serious political analysis. And since the filter is applied with Anglo-Saxon criteria, idiomatically and culturally, the nuances of other languages and other traditions are ground up by filters designed for English.
The criterion of what's considered toxic isn't technical. It's political in the strict sense, in the sense of Carl Schmitt even if almost no data engineer has read Schmitt. It's decided by whoever writes the blocklist. And it's decided, for management economy, without consulting the communities whose cultural production is going to be excluded.
It's not a critique from an automatic woke reading, which the blog avoids on principle. It's an operational observation. Filtering affects the model's statistical base, and the statistical base decides what later comes out of the model's mouth. Whoever filters opines, in silence, on what's sayable and what isn't, and their opinion is applied to the voice of the intelligence that two hundred countries will use to answer their citizens.
The consequence, where it lands
For the Spanish, Italian, Brazilian, Indonesian or Kenyan user, the operational consequence is fairly uniform. The model's answers reach them as if they had been thought in English and mentally translated. The conceptual frames are Anglo-Saxon by default: the "political debate" that appears has the shape of US polarisation, the "family values" carry the nuance of Anglo evangelical Christianity, the "labour disputes" assume the North American individual contract, the "education systems" of reference are K-12 and Ivy League universities. When you ask the model to adapt to the local context, it does, but it starts from a mould that isn't local, and the translation culturally never quite completes.
This doesn't mean the model actively discriminates. It means its weighted median is the corpus's, and the corpus's is that of the Anglo-Saxon digital world hegemonic since the nineties. Asking the model to opine as if it weren't that world doesn't cancel the composition. It adapts it superficially. Underneath, the model's weights are still what they are.
Bender and others, in On the Dangers of Stochastic Parrots (FAccT 2021), drew attention to this point before the product's mass adoption made evident what it implied. When an entire community shares the same assistant, and the assistant thinks by default in one language and one culture, the pressure toward that language and that culture becomes infrastructural. It doesn't act as conscious persuasion. It acts as background.
Safiya Umoja Noble, in Algorithms of Oppression (NYU Press, 2018), had documented the analogous mechanism in search engines. What appears in the first results isn't the product of an editorial decision: it's the product of a set of technical decisions that, aggregated, configure a bias. Her reading was specifically racial. The structural argument is still applicable to LLM territory: the composition of the corpus produces an aggregate bias that acts without anyone having had to write the line of code that declares it.
Who weighs, a better question than what it answers
The serious political question about today's AI isn't what a model answers in a particular debate. That question enters the public debate easily, gives headlines, produces cyclical outrages, and the industry knows how to manage it with post-training adjustments and output filters. The serious question, the hard one, is who weighs.
Who decided the corpus would have that linguistic composition. Who decided which sources to filter. Who decided which digitised libraries were incorporated and which were left out. Who decided which definition of "toxicity" would be used to prune. Who decided that the benchmarks ended up inside the corpus and no one did anything to prevent it. The answer, almost always, is a small team of engineers in a specific company, in a specific country, with criteria that aren't published and aren't submitted to external audit.
Calling them politicians would be unfair in the colloquial sense: they don't see themselves as politicians, they don't run for office, they don't answer to an electorate. Calling them politicians is exact in a broader sense of the term: the one exercised, even if it doesn't name itself that way, by whoever holds the capacity to configure the frame in which others will later take their decisions.
That the choice of dataset remains a technical discussion, opaque, dispersed among companies with protected trade secrets, is the biggest quiet political ceding of the first quarter of the twenty-first century. There are no demonstrations over the composition of the corpus. There's no parliamentary debate over the blocklist. There are no citizen observatories on the linguistic representation in the models entering public use. Everything else gets discussed. That doesn't.
Definitions
Dataset / corpus. A structured set of texts a language model is trained on. Its composition determines, in good part, the statistical base with which the model will later answer.
Common Crawl. A non-profit foundation that has crawled the public web since 2008 and publishes monthly archives of the content gathered. Its archives are the base, direct or filtered, of most open corpora.
C4 (Colossal Clean Crawled Corpus). A filtered, cleaned version of Common Crawl, used by Google's T5 family and derivatives. Documented by Dodge and others (2021).
The Pile. An open 800GB corpus, made up of twenty-two documented sources (Wikipedia, GitHub, books, ArXiv, etc.). Published by EleutherAI in 2020.
Blocklist. A list of terms whose appearance in a document triggers the filtering of the whole document. The main content-filtering tool in web-scale corpora.
Dataset cartography. A metaphor proposed by Kate Crawford to underline that the composition of a corpus reflects decisions of inclusion, scale and centre analogous to those of a map, not neutral properties of reality.
Benchmark contamination. The presence of evaluation examples in the training corpus. It artificially inflates the model's scores on those benchmarks. Documented as a systemic problem since 2021 (see C4 in Dodge et al.).
References
Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. The general framework on the dataset as cartography and on the political materiality of AI.
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. DOI: . A diagnosis of how the composition of the corpus carries over to the model and its outputs.
Noble, S. U. (2018). Algorithms of Oppression. How Search Engines Reinforce Racism. NYU Press. Documentation of aggregate bias in algorithmic systems, transferable to LLM territory.
Gao, L., Biderman, S. et al. (2020). The Pile. An 800GB Dataset of Diverse Text for Language Modeling. arXiv 2101.00027. Detailed documentation of corpus composition, an obligatory reference for understanding what goes in.
Dodge, J., Sap, M., Marasović, A. et al. (2021). Documenting Large Webtext Corpora. A Case Study on the Colossal Clean Crawled Corpus. EMNLP 2021. arXiv 2104.08758. An audit of C4: disproportionate filtering of minorities, presence of machine-generated text, benchmark contamination.
Longpre, S. et al. (2024). The Data Provenance Initiative. A Large Scale Audit of Dataset Licensing and Attribution in AI. Nature Machine Intelligence 6, 975–987. arXiv 2310.16787. An audit of more than eighteen hundred text datasets: licence omission above seventy percent and attribution or categorisation errors above fifty percent in the popular repositories.
Common Crawl Foundation. Statistics of Common Crawl Monthly Archives (language statistics). . Primary data on the linguistic and temporal composition of the archive; the percentages cited correspond to a recent monthly archive and vary between archives.
You might also like
- Invisible biases. The declared bias gets filtered, the baseline one stays intact
- The illusion of neutrality. A neutral AI is an AI whose bias matches yours
- Kate Crawford, the Atlas of AI and what the press has told
- Invisible AI. The one that decides about you without you knowing the decision ever existed

Comments0
No comments yet.
Leave a comment