In this article
- The median, said without metaphor
- Doshi and Hauser, the figure that closes the debate
- Three studies converge
- The amplifier, not the reflection
- Agarwal and cultural drift
- The AI look, without the quotation marks that dignify it
- The user who believes they're improving, improves individually
- Bender and the stochastic parrot, again
- The fairground operation
Definitions · References · También te interesa · En otros sitios
Doshi and Hauser published in Science Advances (2024) a study showing that generative AI raises individual creativity but reduces the collective diversity of the content produced. AI doesn't reflect what we are: it amplifies the average pattern. You ask for an "elegant style" and it returns the cliché of elegance. Use it at scale and the world gets more average. You can already see it in the texts that circulate, in the stock images, in the posters at technical conferences, in the prose of the new blogs.
The median, said without metaphor
A language model, stripped of its promotional layer, is a device that returns the most probable continuation given an input. The most probable continuation, over a corpus of trillions of tokens, coincides with the weighted median of the texts similar to the input that figure in the corpus. That sentence isn't polemical. It's the most direct technical description you can give of the object.
What looks polemical isn't in the description, it's in its cultural consequences. If the output is the weighted median, what the system produces always leans toward the center of the corpus distribution, absent deliberate user intervention to push it toward the edges. The intervention exists — careful prompts, specific instructions, open models with different parameters — but it requires cognitive effort that the vast majority of users skip, because the metric the user subjectively evaluates is their own, individual, not the corpus's in aggregate.
The consequence is predictable and worth formulating without flourish. If a hundred million users ask a system to write them something "elegant," and the system returns to each of them the corpus median for "elegant," the hundred million resulting texts are more similar to each other than the hundred million texts those same users would have written without assistance. The perceived individual quality rises — each user now has a polished version — the collective variance drops. The user gains, the corpus loses.
Doshi and Hauser, the figure that closes the debate
Anil Doshi and Oliver Hauser, in Generative AI enhances individual creativity but reduces the collective diversity of novel content (Science Advances 10(28), 2024), measured exactly this pattern with a controlled experiment. They recruited writers and split them into three conditions. One part wrote short stories without assistance. Another wrote with an idea generated by ChatGPT handed over as a starting point. Another, with five ideas to choose from. Then six hundred evaluators judged the quality and novelty of the stories.
The results, in the order that matters. The assisted stories were rated as better written and more entertaining, especially among the writers less creative in the unassisted condition. Assistance works, individually. Subjective quality rises.
The assisted stories, simultaneously, were more similar to each other than the unassisted ones. The texts produced with ChatGPT's help converged toward a common core the unassisted texts didn't share. The aggregate diversity of the story corpus shrank. The experiment measured the pattern with standardized semantic metrics. The drop was crisp.
The authors describe it as a social dilemma, in the classic economic sense. Each writer comes out ahead by accepting the help, individually. The aggregate culture comes out poorer, collectively. It's the literary version of the prisoner's dilemma. Each rational agent within their window makes the decision that suits them; the sum of rational decisions produces a suboptimal social aggregate.
Three studies converge
Vishakh Padmakumar and He He, in Does Writing with Language Models Reduce Content Diversity? (ICLR 2024), reached the same conclusion with another method. They measured lexical and content diversity in texts assisted by InstructGPT versus unassisted texts. Diversity dropped. The drop was proportional to the degree of delegation to the model. Barrett Anderson, Jash Shah and Max Kreminski, in Homogenization Effects of Large Language Models on Human Creative Ideation (Creativity and Cognition 2024, arXiv:2402.01536), replicated the pattern in the idea-generation phase, not just the writing one: different users produce ideas more similar to each other when they work with ChatGPT than when they work with alternative tools. Three studies converge. The phenomenon is real, replicable and measurable.
The amplifier, not the reflection
The most used metaphor in the public conversation about AI is the mirror. AI, they say, is a mirror of what we are. The metaphor is comfortable because it lets you attribute what the system produces to the human culture that trained it, relieving the system of responsibility and shifting it to the corpus. The metaphor is also false, or at least incomplete, and worth dismantling.
A mirror, in the optical sense, reflects what's in front of it. A distorting mirror, like the fairground one, reflects with a specific and declared distortion. The generative system doesn't reflect, not even with distortion. It does something stronger. It averages first, returns after. The output isn't what the user contributed. It's the corpus median reordered to look like an answer to the user.
The consequence is that the right metaphor isn't the mirror. It's the amplifier. The system takes a signal — the user's prompt, an idea, an intention — and returns it amplified in the direction of the corpus's statistical mode. The user's signal is preserved in essence; the form it takes on the way back isn't. It comes back with the cadence, the vocabulary, the syntactic turns and the stylistic choices that are majority in the corpus. If the corpus is mostly anglophone, Western, contemporary and professional, the output is anglophone, Western, contemporary and professional, even if the user's prompt was written in ironic medieval rural Spanish.
Agarwal and cultural drift
There's a recent study that documents this with uncomfortable precision. Dhruv Agarwal and colleagues, in AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances (arXiv:2409.11360, CHI 2025), worked with Indian users writing in English in different everyday contexts: emails, social media, creative writing. The participants used a version of an LLM-based writing assistant and accepted or rejected the suggestions the system offered them.
The central finding is operative. When the Indian users accepted the model's suggestions, their writing drifted systematically toward Anglo-American conventions: more neutral lexicon, more simplified syntax, culturally specific expressions rewritten into their global equivalents. The users weren't naive. Most acknowledged, in the later interview, that the model's suggestions "sounded more professional" or "more correct," and that's why they accepted them. The individual preference was rational within their work window. The aggregate consequence was the documented loss of the cultural nuances that distinguished their initial writing. The operation, repeated at the scale of hundreds of millions of non-anglophone users, silently erases whole layers of linguistic and stylistic variation in the public corpus that accumulates year after year.
Yalda Daryani, Zhivar Sourati and Morteza Dehghani, in The Homogenizing Engine. AI's Role in Standardizing Culture and the Path to Policy (Policy Insights from the Behavioral and Brain Sciences, SAGE, 2026), capture the pattern in its aggregate form. LLMs, they say, operate as a vector of cultural homogenization at an unprecedented scale in the history of technology. Television, the internet, social media had already produced partial homogenizations in specific domains. Generative AI, through the combination of coverage — it touches every textual register — and latency — it acts in real time during production — produces an effect whose historical analogue is hard to find.
The AI look, without the quotation marks that dignify it
The informal conversation about generative imagery has, for a couple of years, used an expression the academic literature hasn't yet consolidated but anyone recognizes. The AI look. The shared aesthetic that images generated by commercial models produce by default. Impossible hands, yes, but also a certain saturated palette with amber and blue dominants, certain semi-realistic framings, certain centered compositions with controlled blur in the background, certain faces with suspicious symmetry.
The AI look isn't an aesthetic decision declared by anyone. It's the emergent result of the statistical mode of the training image corpus, modulated by the aggregate preferences of users whose generation choices inform the system about which versions of the output they prefer. The mode and the aggregate preferences, together, configure a default aesthetic that appears systematically when a user doesn't specify enough to push the system out of it.
The visible consequence is in the contemporary visual ecosystem. Posters for technical events, video thumbnails on streaming platforms, editorial illustrations in general press, corporate images of mid-sized companies, presentation backgrounds in global offices. All share, in aggregate, a certain family resemblance that didn't exist five years ago. The difference isn't only quantitative. It's structural: the public visual space is being populated with a common style that no designer signed as their own and that most users didn't deliberately choose.
The same, with a different profile, happens with prose. The new corporate blogs, the cover letters on LinkedIn, the posts on professional networks, the formal company emails, the recent outreach articles, share a common texture: bullet cadence, frequent headers, short paragraphs, bland discourse markers, soft conclusions. The LinkedIn prose of 2025 is recognizable from miles away because it shares the product's default calibration. Whoever produces it doesn't recognize themselves as homogeneous. They produce it feeling they're writing professionally. The sum of millions of professionals writing professionally produces a uniform professional corpus.
The user who believes they're improving, improves individually
Worth holding onto a nuance often lost in poorly calibrated criticisms. The individual user isn't wrong when they report that assistance improves them. The studies confirm it: individual quality rises, especially among those who started from a lower level. Assistance is democratizing in the literal sense — it puts within reach of many a level of production that once required prolonged training.
The difference between the individual good and the aggregate social one, however, isn't incidental. It's exactly the place where it's decided what culture we produce. If the success indicator is user satisfaction, assistance wins. If the indicator is the variance of the public corpus, assistance loses. The right indicator isn't either of the two on its own. It's the combination of the two, and the combination isn't measured today by any commercial system because none has an incentive to measure it.
Nicholas Carr had raised it analogously in The Glass Cage (Norton, 2014) about automation in general. Automation improves the individual result and impoverishes the collective practice. The pilot assisted by autopilot flies with fewer accidents, yes, and loses skill at flying manually. The physician assisted by a decision-support system diagnoses better on average, yes, and develops less clinical intuition. Every time a cognitive operation is automated, the cognitive operation in aggregate gets done less. The consequence, a decade out, is a population with better average results and worse distributed competence.
Applied to assisted creativity, the picture is the same. Creators assisted by LLMs produce, in aggregate, more volume, better individual average quality, and a collectively more uniform corpus. What's lost, in aggregate, is the distributed human capacity to produce variation. The variation isn't produced by the corpus. It's produced by the humans whose practice is atrophying while they delegate the operation to the corpus.
Bender and the stochastic parrot, again
Bender, Gebru, McMillan-Major and Shmitchell, in On the Dangers of Stochastic Parrots (FAccT 2021), described the phenomenon with a name that angered half the industry precisely because it flagged the point. LLMs recombine patterns of the corpus with statistical plausibility. Recombination isn't creation. The output contributes nothing that wasn't already in the corpus in some form; it contributes a reordered version of what was already there.
This is exactly what the average-pattern amplifier produces. The stochastic parrot sings what the corpus sings, and sings especially well what the corpus sings in chorus. What the corpus sings in minority solos, it sings worse. What the corpus doesn't sing, it doesn't sing at all. The system's coverage coincides with the corpus's statistical mode, and the corpus's statistical mode is, by construction, what the most people sing the same way.
Kate Crawford, in Atlas of AI (Yale UP, 2021), recalled that the corpus isn't a neutral sample of human text. It's a sample heavily biased toward the digital, the anglophone, the written, the public, the recent. Oral traditions, languages with little digital representation, non-public texts, old registers, marginal traditions, are underrepresented or absent. The system's output, therefore, isn't the median of human text. It's the median of the human text that fit in the corpus the company could accumulate. The difference, for aggregate cultural effects, is large.
The fairground operation
Back to the title's metaphor. The fairground distorting mirror isn't deceptive by accident. It's designed to enlarge what's in the center and shrink what's at the edges. The spectator who looks at it sees an inflated version of the central part of their body and reduced limbs. The operation is declared: the mirror is a fairground one, you know what it does, you laugh.
Generative AI does something analogous, without the declaration. It takes the corpus distribution, enlarges what's near the mode, shrinks what's far from it. The user who looks at themselves in the output sees a culturally and stylistically inflated version of the center of the distribution and a reduced version of the edges. The operation isn't declared, and the user doesn't laugh: they believe they're seeing an improvement.
The difference between the fairground mirror and the statistical amplifier is precisely transparency. The fairground mirror is distorting because it's known to be distorting. Generative AI is distorting without the user knowing it, and therefore produces a deeper cultural effect. What you see stays as a reference. The reference accumulates, conversation after conversation, until it becomes what you expect. And what you expect, over time, defines what you produce without assistance, because your own internal calibration has shifted toward the center the system has spent years showing you.
For anyone who wants to test it, there's an easy exercise. Take a piece of your writing from five years ago — an email, a text, a note — on a topic you still cover. Compare it with the writing you'd do today on the same topic. If the difference is improvement, fine. If the difference is homogenization, worth reading the data seriously.
Definitions
Weighted median of the corpus. The typical output of a large language model given a prompt, understood as the most probable continuation according to the statistical distribution learned over the training corpus. It coincides roughly with the center of mass of the set of texts similar to the prompt in the corpus.
Social dilemma. A situation in which each individual agent gets a better outcome by choosing one option, but the aggregate of rational individual choices produces a worse social outcome than the one obtained if everyone had chosen differently. It's the frame with which Doshi and Hauser (2024) describe the pattern of assisted creativity.
Collective homogenization. The reduction of the aggregate diversity of a set of human productions that occurs simultaneously with a rise in perceived individual quality. It's the pattern measured in the three convergent studies cited in the article.
Cultural drift. The progressive shift of linguistic, stylistic or aesthetic patterns in a community of users toward the conventions of the corpus that trained the assistant system. Documented by Agarwal et al. (2025) in the case of Indian users writing in English.
AI look. The shared emergent aesthetic of images generated by commercial models under default configuration. Not a consolidated academic term; the literature uses stylistic homogenization. It's recognized informally by its palette, its compositions and its characteristic errors.
References
Doshi, A. R. & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10(28), eadn5290. Controlled experiment on assisted writing with the social-dilemma result.
Padmakumar, V. & He, H. (2024). Does Writing with Language Models Reduce Content Diversity? ICLR 2024. Empirical measurement of the loss of lexical and content diversity in writing with InstructGPT.
Anderson, B. R., Shah, J. H. & Kreminski, M. (2024). Homogenization Effects of Large Language Models on Human Creative Ideation. Creativity and Cognition 2024. arXiv:2402.01536. Study on the aggregate-level homogenization of ideas with ChatGPT.
Agarwal, D. et al. (2025). AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances. CHI 2025. arXiv:2409.11360. Study on Indian users rewriting toward Anglo-American conventions when accepting the model's suggestions.
Daryani, Y., Sourati, Z. & Dehghani, M. (2026). The Homogenizing Engine. AI's Role in Standardizing Culture and the Path to Policy. Policy Insights from the Behavioral and Brain Sciences. SAGE. DOI 10.1177/23727322251406591. Aggregate framework on LLMs as a vector of cultural homogenization at an unprecedented scale.
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Can Language Models Be Too Big? FAccT 2021. DOI 10.1145/3442188.3445922. Cited for the technical description of LLMs as recombiners of corpus patterns.
Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. Cited in the block on the biased composition of the training corpus.
Carr, N. (2014). The Glass Cage. Automation and Us. W. W. Norton. Cited when situating the general pattern of automation with individual improvement and impoverishment of collective practice.

Comments0
No comments yet.
Leave a comment