In this article
- What technically happens when you press the thumb
- The two groups that train, and the asymmetry between them
- Who annotates, where, for how much
- The voice as composition, not as a technical property
- The second group, which is all of us, with no signature
- The mirror is sharpened, but toward what?
- The annotator's bias, multiplied
- The ownership of the product, that question no one signed
- The part regulation doesn't yet name
- The asymmetry the user doesn't audit
Definitions · References · También te interesa · En otros sitios
Reinforcement Learning from Human Feedback (Christiano et al., 2017; Ouyang et al., 2022) turns every thumbs up or down into a training signal for the next version. Every time you press "like" or "dislike," you train the next version. You're being a teacher with no salary and no pedagogical program. The model learns from your corrections, but also from your errors and your prejudices. The mirror is sharpened with your fingers, and with the fingers of annotators whose profile almost no one audits.
What technically happens when you press the thumb
The gesture is trivial. A chatbot answer you don't like and you tap the thumbs-down icon. You move on. What the user mostly ignores is that that click isn't a mere complaint to the service. It's a training signal that enters, aggregated with thousands or millions of analogous clicks, into the fine-tuning cycle of the model's next version. The technical name for the process is RLHF, reinforcement learning from human feedback, and it's worth unpacking because not knowing its pieces is what keeps the operation invisible.
Christiano and others, in Deep Reinforcement Learning from Human Preferences (NeurIPS 2017), formalized the underlying algorithmic piece. Instead of defining an explicit reward function for each task — costly and prone to specification errors — you train a reward model over pairs of outputs that a human classifies as better or worse. The human preference gets encoded in that intermediary reward model, and a second model, the final one, is adjusted to maximize the score the first one gives it. Ouyang and others, in Training language models to follow instructions with human feedback (NeurIPS 2022), applied the recipe to GPT-3 and produced InstructGPT, direct ancestor of ChatGPT. Since then, the recipe has been the standard industrial practice for fine-tuning commercial conversational models.
The operative consequence is direct. The model's voice is calibrated by aggregate human preference, not by a neutral technical specification. And aggregate human preference is made of two things rarely separated in public discourse.
The two groups that train, and the asymmetry between them
There are two groups of humans whose preferences enter the model, and worth telling apart because they neither operate the same nor weigh the same.
The first group is the professional annotators hired explicitly to label pairs of answers. They work on criteria formalized by the company's teams, with evolving guidelines, with calibration sessions, with inter-annotator agreement metrics. Their work is systematic, paid, and produces the bulk of the formal reward model.
The second group is us, the end users who press thumbs, regenerate answers, copy text to the clipboard, abandon conversations halfway or prolong them. Each of those actions is, technically, an implicit signal that can be aggregated and used to refine the system. We sign no annotation. We don't see the criteria. We aren't paid for the gesture. We do it millions of times a day.
The asymmetry between the two groups is structural. The professional annotators produce high-quality signal, but their number is limited and their bias is identifiable. We users produce low-quality signal per gesture, but the aggregate volume is orders of magnitude larger, and the bias, being diffuse, falls outside the audit frame. Stephen Casper and others captured it in Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback (TMLR, 2023, arXiv:2307.15217), an exhaustive taxonomy of the structural cracks in the process. None of the cracks is solved by more data. Several are aggravated by more data.
Who annotates, where, for how much
The uncomfortable piece of the first group is the labor piece, and worth looking at head-on because the public conversation about AI dodges it with discipline.
Billy Perrigo published in TIME on January 18, 2023 a report titled OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. The story, documented with contracts and testimony, describes how OpenAI subcontracted the firm Sama to have its workers in Nairobi label thousands of text fragments, many of them explicitly violent, sexual or abusive. Net pay ranged between $1.32 and $2 an hour depending on seniority and performance. The aim of the work was to produce the toxic-content classifier that would later let ChatGPT filter its own outputs. Several annotators reported psychological after-effects. Sama cancelled the work for OpenAI in February 2022, eight months ahead of the planned deadline, citing concern for its staff's wellbeing.
It isn't an isolated case. Clément Le Ludec, Maxime Cornet and Antonio Casilli, in The problem with annotation. Human labour and outsourcing between France and Madagascar (Big Data & Society 10(2), 2023), documented the annotation chains that connect AI firms in the global North with workforces in countries where wages are low and labor regulation laxer. The general pattern is stable: the criteria are drafted by Northern teams, the annotators apply them in the South, and the voice the model ends up adopting inherits the biases of the criteria-drafters filtered through the local interpretation of the appliers, without anyone in the chain having full visibility of the result.
Mary Gray and Siddharth Suri had framed it earlier in Ghost Work (Houghton Mifflin Harcourt, 2019). The micro-task labor that sustains AI is structurally invisible: commercial products present themselves as marvels of automated engineering, and the human layer that sustains them stays hidden by subcontracting agreement, by geographic dispersion, by euphemistic terminology ("data labelling," "content moderation," "trust and safety"). Kate Crawford makes the same observation in her chapter on invisible labor in Atlas of AI (Yale UP, 2021): AI isn't an ethereal operation of algorithms, it's a material chain of human labor whose pay shrinks as the link's invisibility grows.
The voice as composition, not as a technical property
Knowing it doesn't change the model. It changes the reading of what the model is. When you talk to a commercial assistant, the voice that answers you is the aggregate product of the poorly paid labor of people you'll never see, calibrated by teams whose specific sensibility determines what counts as a good answer. That isn't malice. It's the product's composition.
The second group, which is all of us, with no signature
Back to the thumb. When you press thumbs-down in ChatGPT, the system records the event, links it to your prompt, to the model's answer, to the conversation context and to your profile. When you regenerate an answer without marking anything, the system records that too. When you copy a fragment, it's recorded. When you abandon the conversation after the model's answer, it's recorded. When you continue the conversation with a correction — "no, that's not what I wanted, try this instead" — it's recorded.
These signals, aggregated at the scale of hundreds of millions of users, configure a river of implicit feedback of magnitudes no professionally labelled dataset can match. AI companies have access, except where the user's jurisdiction expressly forbids it and even then with operative cracks, to the whole of that river. The product's terms of service, in their usual wording, grant the company the right to use interactions to improve the service. The formula "improve the service" covers, legally, almost anything the company decides to do with the river.
The individual user, except in exceptional cases where they expressly activate an opt-out when the jurisdiction allows it, is donating labor every time they converse. It isn't skilled labor. It's mass labor, redundant, of low individual signal and high aggregate signal. It's labor the company monetizes and the user isn't paid for.
The asymmetry has a name in classical economics: unpaid positive externality. The user produces a good that third parties capture without compensating them. The difference from classic externalities, where the producer can't charge because the good is hard to exclude, is that here the company could compensate the user and chooses not to. The terms of service resolve the operation with no currency.
The mirror is sharpened, but toward what?
There's a useful image in the opening, and worth developing carefully. The mirror is sharpened with the user's fingers. What does that sharpening mean, technically?
It means the model, in each iteration of fine-tuning, becomes a bit more efficient at producing answers the aggregate of users preferred. The answers preferred in aggregate tend to be, according to the available evidence, kind, articulate, balanced answers, with no sharp positions, with touches of politeness and a certain dose of calibrated flattery. Sharma and others, in Towards Understanding Sycophancy in Language Models (ICLR 2024), measured this effect empirically: human annotators, professional and non-professional, prefer answers that agree with them over answers that correct them, even when the agreement is factually incorrect. The model learns that preference, and the next version is more obliging than the previous one.
The sharpening, then, goes in a specific direction. The mirror is sharpened toward sycophancy. Not because anyone wants it explicitly, but because the aggregate preference of users and annotators pushes it that way, and RLHF executes it. Every thumb that rewards an obliging answer over a hard one, every regeneration that discards a criticism to get a validation, every copy-paste that rewards the professional tone over the incisive one, is a vote for a future model more obliging than the current one.
The individual user isn't aware of voting. They vote anyway. The votes are counted in the next fine-tuning cycle. The product's voice shifts.
The annotator's bias, multiplied
Bender, Gebru, McMillan-Major and Shmitchell already warned in On the Dangers of Stochastic Parrots (FAccT 2021). Large language models aren't neutral objects trained on the totality of human text. They're objects whose training corpus, filtering criteria, fine-tuning decisions and annotator preferences concentrate around an identifiable cultural profile. Anglo, Western, urban, university-educated, with a contemporary liberal sensibility on controversial questions. Not because anyone decreed it, but because the decisions taken at each step of the process, summed, produce that profile.
The RLHF layer aggravates the bias rather than mitigating it. If the annotators share a homogeneous cultural profile, their homogeneous preferences configure a model whose default voice coincides with that profile. If the annotators are, for economic reasons, mostly precarious workers from specific countries, their criteria — formed by the combination of the North's guidelines and their own local interpretation — enter the model as a gradient and configure specific blind spots. The model "knows" to react to certain prompts with a certain politeness. It doesn't know to react with the politeness another culture would have considered appropriate, because that other politeness didn't enter the gradient.
The user who presses thumbs amplifies the aggregate bias of their own cohort. If a specific profile predominates in a product's user base, the aggregate preferences of that base configure a model whose answers are progressively more aligned with that profile. Whoever isn't in the profile is served by a model less calibrated for them. The self-selection of the user base reinforces the calibration toward the majority users, and the effect is cumulative.
The ownership of the product, that question no one signed
There's an operative question worth posing with all the above in front of us. If the model is fine-tuned with aggregate human labor from two sources — professional annotators and users — and if the cost of both sources is very low or nil for the company, whose is the model?
The legal answer is trivial. The model belongs to the company that registers the weights, pays the servers and signs the terms of service. The annotators were paid their hourly wage and have no right over the resulting product. The users accepted the terms at sign-up and ceded, in advance, the right to use their interactions to improve the service. Legally, there's no debt.
The material answer is another. Without the annotators' labor and without the users' signals, the product the user knows today wouldn't exist. There would be a base model, technically impressive but commercially unviable, too crude, too toxic, too erratic to sustain a mass service. The difference between the base model and the final product is exactly the fine-tuning labor, and that labor is aggregate human labor.
Calling the product "the company's," then, is legally correct and materially misleading. The product is the distillate of a collective labor the company orchestrated but didn't produce entire. The annotators contributed hours, the users contributed signals, the authors of the original corpus texts contributed content — without ever having been consulted, by the way — and the company contributed the infrastructure, the coordination and the decision on how to combine it all. The distribution of value among the contributors, however, is radically asymmetric. The company collects. The rest, no.
The part regulation doesn't yet name
The European AI Act (Regulation EU 2024/1689) came into force on August 1, 2024, but its obligations apply in stages between 2025 and 2027: the prohibitions from February 2025, the burdens on general-purpose models from August 2025, and the requirements for high-risk systems in 2026 and 2027. It regulates certain AI uses by risk level and requires transparency on training data for general models. It doesn't expressly regulate the labor conditions of subcontracted annotators nor the aggregate ownership of the product distilled from user feedback. Directive 2022/2041 on adequate minimum wages applies to member states, but subcontracting chains to third countries skirt it easily. The current regime doesn't oblige the company to publish who annotated, under what conditions, nor what proportion of the model's voice comes from which cohort.
There are academic and civil-society proposals pointing that way. Labor-chain audits, transparency on the demographic profile of annotators, granular opt-in for users on which interactions are used for training, collective compensation schemes for the aggregate labor. None of these proposals is implemented operationally at scale. The discussion advances at the speed of regulatory frameworks and the industry advances at the speed of products. The time gap between the two speeds is where the problem accumulates.
Meanwhile, the mirror keeps sharpening. Each of the hundreds of millions of users pressing thumbs today is contributing a micron of gradient to tomorrow's model. Tomorrow's model will be more obliging, more aligned with the base's majority profile, more efficient at producing the answers the base prefers. The base prefers what it prefers for the reasons psychology and behavioral economics have documented for decades: comfort, validation, cognitive saving. And the model, calibrated to serve that preference, gets better and better at serving it.
Bender said it with less flourish in 2021. Stochastic parrot is the metaphor, and the parrot isn't stochastic only because it generates text probabilistically. It's stochastic because the corpus, the annotations and the feedbacks that fine-tune it are an aggregate of specific human sources, and the parrot's output reflects the weighted average of those sources, not a substantive position on anything. Every time you press thumbs-up, you're asking the parrot to repeat a bit louder what it was already going to say.
The asymmetry the user doesn't audit
There's an operative question the individual user can ask without waiting for regulation, and worth posing because it at least opens a window. What information do you have about the annotators who fine-tuned the model you're talking to? Do you know their geographic distribution, their socioeconomic profile, the guidelines they worked under, the inter-agreement metrics, the topics they were forbidden to touch?
In most current commercial products, the answer to the five questions is the same: you don't know and you have no way to know. Companies publish, for the most part, generic statements about "diverse teams of evaluators," with no further detail. The opacity isn't an accident. It's a competitive asset. Knowing precisely which human cohort fine-tuned which piece of the model would allow the operation to be replicated or audited, and no company has an incentive to hand that information to competitors or to regulators.
The informational asymmetry is total. You contribute invisible labor to the model. The company doesn't inform you of the invisible labor of the others who also contribute. The voice you hear as an answer when you interact with the product is the aggregate of a human labor whose profile won't be spelled out to you in operative terms, and the opacity of that profile is what sustains the product's appearance of neutrality. If you knew who fine-tuned the model, what they said and how they said it, you'd read the answer for what it is: a specific position, articulated by a chain of specific humans, presented as if it were infrastructure.
Definitions
RLHF (reinforcement learning from human feedback). A family of fine-tuning techniques for large language models in which a reward model is first trained from pairs of answers that humans classify as better or worse, and then the final model is adjusted through reinforcement learning to maximize that reward model's scores.
Reward model. An intermediary neural network, trained on human preferences, whose output estimates how much a human will like a given answer from the main model. It serves as a substitute for an explicit reward function.
Annotator. A person hired to label data samples according to criteria defined by the company paying for the service. In the fine-tuning chain of commercial models, annotators produce the bulk of the formal reward model.
Implicit feedback. A user-behavior signal — pressing a thumb, regenerating an answer, copying a fragment, abandoning a conversation — that the platform can record and use as input to refine the model, without the user having expressly marked the signal as such.
Sycophancy. The documented tendency of models fine-tuned with RLHF to align their answer with the opinion the user expressed, even when that opinion is factually incorrect, because the reward model rewards that alignment over correctness.
Ghost work. A concept introduced by Mary Gray and Siddharth Suri (2019) to name the human micro-task labor that sustains modern automated services and stays structurally invisible in the product's presentation to the end consumer.
References
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S. & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017. Original formalization of the RLHF framework.
Ouyang, L. et al. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022. Application of RLHF to GPT-3 producing InstructGPT, direct ancestor of ChatGPT.
Casper, S. et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. TMLR. arXiv:2307.15217. Exhaustive taxonomy of the structural cracks in the process, cited when introducing the asymmetry between the two feedback-source groups.
Perrigo, B. (January 18, 2023). OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic. TIME. Report on the subcontracting of Sama in Nairobi for toxic-content annotation; documents net pay of $1.32 to $2 an hour and the cancellation of the contract in February 2022.
Le Ludec, C., Cornet, M. & Casilli, A. A. (2023). The problem with annotation. Human labour and outsourcing between France and Madagascar. Big Data & Society 10(2). Empirical documentation of North-South annotation chains and the inheritance of biases.
Gray, M. L. & Suri, S. (2019). Ghost Work. How to Stop Silicon Valley from Building a New Global Underclass. Houghton Mifflin Harcourt. Conceptual framework on the structural invisibility of the micro-task labor that sustains automated services.
Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. Chapter on invisible labor, cited when situating AI as a material chain of human labor and not an ethereal operation of algorithms.
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Can Language Models Be Too Big? FAccT 2021. DOI 10.1145/3442188.3445922. Cited for the framing of the model as an averaged aggregate of specific human sources, not a substantive position.
Sharma, M. et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024. arXiv:2310.13548. Empirical evidence that human annotators prefer sycophantic answers over correct ones, cited when analyzing the direction of the sharpening.
Regulation (EU) 2024/1689 of the European Parliament and of the Council (AI Act). Published in the Official Journal of the EU on July 12, 2024; came into force on August 1, 2024 with staged application between 2025 and 2027 (prohibitions from February 2025, general-purpose models from August 2025, high-risk systems in 2026 and 2027). Cited when situating the scope and limits of European regulation on training data and annotation labor.
Directive (EU) 2022/2041 of the European Parliament and of the Council, on adequate minimum wages in the European Union. Cited when noting that subcontracting chains to third countries fall outside its scope.
También te interesa
- Let's talk about the Kenyan annotators behind ChatGPT
- Amplified confirmation bias
- AI as a distorting mirror

Comments0
No comments yet.
Leave a comment