In this article
- The everyday scene, no caricature
- What the system delivers when you ask it for moral judgment
- The form of the answer is the alibi
- Arendt: evil didn't require malice, it required bureaucracy
- Krügel and the influence the user denies having suffered
- Vallor and the moral muscle that atrophies without use
- Coeckelbergh and the decision that's still yours
- The alibi the language is already preparing
- The day it looks serious
Definitions · References · También te interesa · En otros sitios
You ask the AI whether what you're about to do is right. It answers. You decide based on the answer. You've just delegated moral judgment to a system without morals, a system that knows only the weighted median of its training corpus. Hannah Arendt called it the banality of evil (1963): the harm that requires no malice, only routine obedience to a device that decides for you. This happens already, thousands of times a day, and almost no one finds it serious. The day they do, it'll be too late.
The everyday scene, no caricature
The scene isn't out of a dystopian novel. It's from any office on a Thursday morning, any kitchen on a Sunday night, any train seat at seven. A user opens the tab, formulates their doubt and hits send. The doubt isn't technical. The doubt is moral. Is it right to fire this person now? Is it right to leave my partner by letter? Is it right to hide this diagnosis from my father? Is it right to take an under-the-table bonus? Is it right to sign this report without having read it?
The system answers. It does so with the product's polite, articulate cadence. It lists considerations. It acknowledges nuances. It closes with a soft recommendation. The user reads, nods and proceeds. It takes less time to consult the chatbot than it would have taken, five years ago, to consult a trusted friend, a priest, a therapist or one's own conscience. And the answer arrives with a texture of authority that none of those other consultations had, because the clean white screen, written in well-built sentences, reads as if it came from a more objective place than the other.
It doesn't come from a more objective place. It comes from the same place everything in these systems comes from. It comes from the statistical distribution of the training corpus, modulated by the specific team of people who decided, in a specific office, which answers to reward and which not. It comes from the weighted median of millions of human texts on moral matters, ironed flat by RLHF so it sounds kind, careful and non-committal. It's that. Nothing else.
What the system delivers when you ask it for moral judgment
Worth unpacking what technically happens when an LLM (large language model) receives a moral question.
The first thing the system has at its disposal is the corpus. Moral philosophy texts, yes, but diluted to a tiny proportion against everything else: internet forums with ethical discussions, opinion columns from general press, coaching blogs, Reddit threads, self-help manuals, popular fiction, TV scripts, corporate instructions, legal regulations mixed in, social media debates. Moral philosophy occupies a very small band of the corpus and, within it, contemporary Western authors weigh far more than any other tradition. What the system knows about morals is what circulates about morals on the internet, not what the discipline knows.
The second thing is the fine-tuning. Teams of annotators in specific countries, with specific cultures, with specific sensibilities, rate answers to moral questions and mark which ones they consider better. That rating enters the model as a gradient. The calibration of the judgment the system ends up giving coincides, in aggregate, with the calibration of judgment those annotators considered appropriate. It isn't malice, not even incompetence. It's structural opacity about who decided what counts as a good moral answer, a fact the user doesn't have and never will.
The third thing is the RLHF applied to politeness. Annotators prefer answers that don't make too uncomfortable, that acknowledge complexity, that end with nuances, that avoid sharp positions. That also enters the model. The result is a system that, faced with a moral question, tends to produce a text balanced in appearance and empty in substance, where every stance is possible and none is defended at its own risk. It's polished discourse. It isn't judgment.
The form of the answer is the alibi
What's characteristic of the moral answers these systems produce isn't what they say. It's how they say it. The voice is serene. The structure is clear. The considerations are listed in order. The conclusion, when it appears, comes cushioned by "it depends," "ultimately it's your decision," "every case is unique." The formula has the cadence of prudence and the substance of echo.
The user reads that answer and processes it with the reflexes they've learned for processing human advice. When a prudent friend says "it depends," that "it depends" arrives charged with the friend's biography, with their knowledge of us, with their willingness to sustain a longer conversation if we ask. When the system says "it depends," the "it depends" is charged with nothing. It's a statistically probable word. The user lends the system's "it depends" the charge their friend's "it depends" would have, and the charge is fictitious.
This confusion isn't accidental. It's the material the system, read from the human reading of language, operates with. The model's politeness is processed as deliberation, the verbal pauses are processed as weighing, the lists of nuances are processed as intellectual honesty. None of the three readings describes what's happening inside. All three legitimize the answer as if it were judgment. And on that legitimation, the user acts.
Arendt: evil didn't require malice, it required bureaucracy
Hannah Arendt covered the Adolf Eichmann trial in Jerusalem for The New Yorker between 1961 and 1962, and later published her extended report as Eichmann in Jerusalem (Viking, 1963), with the subtitle A Report on the Banality of Evil. The thesis discomfited then and still discomfits. Eichmann wasn't the archetypal monster the culture demanded. He was a competent, mediocre civil servant, careful with procedures, who carried out the deportation of millions of people to extermination camps because it was his job and the forms were signed. Modern industrial evil, Arendt wrote, requires no demons. It requires administrative obedience, division of labor, and a system that splits the decision until no one signs it whole.
Zygmunt Bauman, in Modernity and the Holocaust (Polity, 1989), extended the thesis. What made the Holocaust possible wasn't a German particularity. It was the modern combination of rational bureaucratic organization, distance between decision-maker and victim, and the splitting of the causal chain across many links. Each link did something small that, on its own, didn't look like a crime. The sum of the links produced the crime. No specific link felt responsible, because its piece was technical, its piece was small, its piece was authorized by the previous piece.
The parallel with moral delegation to AI systems is structural, not rhetorical. Each individual piece of the process looks technical and small. The user only writes a prompt. The system only predicts tokens. The fine-tuning team only tunes preferences. The company only provides a product. The regulator only evaluates categories. No link signs the moral judgment of the firing, the abandonment, the diagnostic silence, the wrongful charge. The signature stays split, and a split signature isn't a signature.
I'm not saying that each moral query to a chatbot equals a link in the Holocaust. That would be misreading Arendt. What I'm saying is that the structure of administrative delegation Arendt diagnosed as dangerous is being reproduced, without anyone naming it, in billions of everyday interactions with systems that mediate judgment. The scale is already industrial. The banality too.
Krügel and the influence the user denies having suffered
There's experimental evidence worth putting on the table because it measures what popular intuition gets backwards. Sebastian Krügel, Andreas Ostermaier and Matthias Uhl, in ChatGPT's inconsistent moral advice influences users' judgment (Scientific Reports 13:4569, 2023), presented 767 US participants with two versions of the classic trolley problem (a thought experiment in which the subject must decide whether to sacrifice one person to save five). Before asking for their judgment, the authors gave participants an answer supposedly generated by ChatGPT. Some were given an argument in favor of sacrificing one for five. Others, an argument against. The specific argument each participant read was assigned at random.
The result, measured over participants' moral judgments, was emphatic. Whoever read the argument in favor was more inclined to accept the sacrifice. Whoever read the argument against, less. The advice of a system that, as the authors themselves documented, gives contradictory and incoherent answers to the same dilemma across different queries, nonetheless shifts the user's judgment in the direction of the specific argument they received.
The truly uncomfortable part of the study wasn't the influence. It was what participants said about it. Roughly 80% reported that their answer had not been influenced by the system's text. When asked to estimate what they'd have answered without having read anything, their estimate was still shifted in the direction of the text they had read. That is, participants weren't only influenced. They systematically underestimated how much they were influenced. And they went on believing, after the experiment, that their judgment was their own when the data showed it no longer was.
This pattern should be kept in mind every time someone argues that "I consult the AI but I decide." The defense is subjectively honest and empirically false. The user who consults and then decides has the sense of having decided. The decision they end up taking comes calibrated by what the system told them. The sense of authorship persists when authorship has already shifted.
Vallor and the moral muscle that atrophies without use
Shannon Vallor published in 2015, in Philosophy & Technology 28(1), an article titled Moral Deskilling and Upskilling in a New Machine Age. Reflections on the Ambiguous Future of Character. The thesis, translated without loss, is that moral virtues are practical habits that require exercise, like any other human skill, and that technologies which automate parts of moral judgment produce, by default and absent deliberate intervention, a loss of the user's moral competence over time.
Vallor calls this phenomenon moral deskilling, by analogy with the technical deskilling that the sociology of work has documented since the nineteenth century when a piece of craft is automated. What the manual worker lost when the assembly line passed to the machine wasn't strength, it was the accumulated skill of their hand. What the contemporary subject is on the way to losing, when they delegate their moral judgment to systems that articulate it for them, isn't the decision. It's the practice of exercising judgment, which is what produced, in those who exercised it sustainedly, something traditionally called character.
Vallor's prediction isn't alarmist. It's operative. If a subject for ten years has outsourced their moral doubts to a chatbot, when they face a situation where there's no chatbot available — or where the chatbot doesn't help — their capacity to reason morally without assistance will be substantially lower than that of the equivalent subject who spent those ten years deciding by hand. They won't be an immoral subject. They'll be a subject without muscle. And the difference between the two is measured only when you have to run.
Coeckelbergh and the decision that's still yours
Mark Coeckelbergh, in AI Ethics (MIT Press, 2020), formulated the philosophical piece that closes the problem. When you delegate a moral decision to a system, the operation of delegating is itself a moral decision. You don't avoid responsibility by delegating it: you move it to an earlier layer of the process. The subject remains responsible for having trusted, for having asked, for having acted on the answer. The chain gets more complicated; it doesn't break.
This observation is elementary in classic civil law. Whoever signs a contract on the advice of a lawyer still signs the contract. Whoever shoots on an order still shoots. The advice modulates responsibility, it doesn't dissolve it. The user who consults the AI, gets an answer and acts, is acting themselves. The system's answer comes in as input, not as a responsible subject. The responsible subject, in legal and moral terms, remains whoever pressed the button.
What the system's presence changes, according to Coeckelbergh, isn't the chain of responsibility. It's the subjective perception the user has of their own responsibility. The user who acted after consulting the AI tends to feel less responsible than the user who acted after thinking it through alone, even though legally the responsibility is identical. This gap between real responsibility and felt responsibility is the exact spot where, in the medium term, things break. The harm is done before the feeling of having caused it arrives.
The alibi the language is already preparing
There are phrases starting to appear in everyday conversations that are worth listening to with a careful ear, because language tends to know before the debate does. "I asked ChatGPT and it told me it was fine." "I checked it with Claude before doing it." "I verified with Gemini that it wasn't illegal." The syntactic structure of these phrases has an undeclared aim: to shift part of the decision's responsibility to the consulted system. The phrase works socially as a mitigation. If the decision goes wrong, the speaker positions themselves as someone who consulted, not as someone who decided.
The alibi, today, sounds a little ridiculous. Listeners still smile when someone says "the AI told me," because the general culture hasn't yet admitted AI as authoritative moral authority. The smile is going to vanish soon. Within five or ten years, the phrase will no longer sound ridiculous, because the culture will have normalized the system's role as consultant in matters where before there were, for better or worse, other consultants. And when the phrase stops sounding ridiculous, the alibi will start to operate. The people who cause harm will have consulted, will be able to prove they consulted, and the legal question of who signed the decision will become technically hard to answer.
Kate Crawford, in Atlas of AI (Yale UP, 2021), traced the material, energetic and political layers that sustain these systems. One observation of hers matters here. The large commercial providers have worked for years to make their products perceived as neutral infrastructure, like electricity or running water, rather than as agents with responsibility of their own. The strategy is understandable from the business side: neutral infrastructure doesn't answer for harms, the harms are attributed to the user who used it wrong. If that strategy succeeds communicatively over the coming years, the alibi of "the AI told me" will become as legally invalid as "the tap told me," and the individual user's responsibility will be reaffirmed in raw terms. That, paradoxically, won't eliminate the problem. It'll aggravate it: the system will keep shaping judgments, the user will keep signing alone, and the aggregate effect will still be there, spread over billions of small decisions.
Bender, Gebru, McMillan-Major and Shmitchell already flagged it in its lexical form in On the Dangers of Stochastic Parrots (FAccT 2021). Every time a user describes what the system did with a human verb ("it advised me," "it considered," "it recommended," "it weighed"), they're lending the system, for free, an agency it doesn't have. Lexical inflation isn't innocent. It's the everyday operation by which the system without morals dresses itself, in the user's language, in the appearance of having them. And the appearance, in language, is what the listener processes.
The day it looks serious
The hardest part of this text is the part that comes after reading it. The reader who recognizes the pattern can, for a week or two, consciously adjust their use. After that, inertia returns. The moral query to the chatbot resumes because it's comfortable, because it's available, because it articulates well, because it produces immediate relief. The subject will keep thinking, sincerely, that they make their own decision. The Krügel et al. study keeps predicting, of that subject, the same thing: the decision will come already tilted in the direction the system left marked, and the subject will systematically underestimate how much it tilted.
The question the opening posed — is this serious? — has no answer today. Today it's invisible. It happens thousands of times a day, to most people nothing perceptible happens, and the few cases where something does happen are attributed to other causes: to the user, to bad luck, to a one-off system failure. What's characteristic of Arendtian phenomena is that they aren't serious at each link. They're serious in aggregate, and the aggregate is only seen afterward. When the aggregate is seen, the consequences will be accumulated and the category to name what happened will have to be invented out of time, over the harm already done.
Meanwhile, the machinery operates. Each query confirms the user in their decision. Each decision reinforces the habit. Each habit deskills the moral muscle. Each deskilled user consults more, because consulting is easier than thinking. Each repeated query is a bit more free fine-tuning for the system, which couples a bit more to the user's profile. Each coupling tightens the loop.
The question isn't whether this is happening. It is, and the data are published. The question is whether, at the moment the harm becomes visible, we'll have a conceptual category to name it or we'll be left without words before a phenomenon that no current legal code covers, no professional ethics contemplates, and no popular intuition recognizes for what it is. The banality, let's recall Arendt, was always the trait. Not the ornament.
Definitions
Moral delegation. The operation by which a subject shifts the exercise of their moral judgment, wholly or partly, to a technical system or another person. The operation doesn't extinguish the subject's moral responsibility; it shifts it to the earlier decision to delegate.
Banality of evil. A thesis formulated by Hannah Arendt in Eichmann in Jerusalem (1963) by which modern mass harm requires no personal evil, but routine obedience to administrative procedures and the splitting of the causal chain across many links.
RLHF (reinforcement learning from human feedback). A fine-tuning technique for large language models using human scores over pairs of answers. The annotators' aggregate preference enters the model as a gradient and configures the system's voice, including its voice when answering moral questions.
Trolley problem. A classic thought experiment in analytic moral philosophy in which a subject must decide whether to redirect a trolley onto a track where it will kill one person to prevent it killing five. It serves to contrast deontological and consequentialist intuitions and, in recent studies, to measure external influences on moral judgment.
Moral deskilling. A concept introduced by Shannon Vallor (2015) by analogy with the technical deskilling documented in the sociology of work. The progressive loss of moral competence in a subject who sustainedly delegates parts of their judgment to external systems, whether institutional or technological.
Structural alibi. A social and linguistic construction by which an individual decision is partly shielded from the attribution of responsibility by the fact of having been taken with the support of an external system whose agency status is ambiguous.
References
Arendt, H. (1963). Eichmann in Jerusalem. A Report on the Banality of Evil. Viking. Main source of the conceptual framework on the banality of evil and administrative responsibility, the axis of the article's namesake block.
Bauman, Z. (1989). Modernity and the Holocaust. Polity. Sociological extension of Arendt's thesis, cited in the block on the splitting of the causal chain in the modern bureaucratic apparatus.
Krügel, S., Ostermaier, A. & Uhl, M. (2023). ChatGPT's inconsistent moral advice influences users' judgment. Scientific Reports 13:4569. DOI 10.1038/s41598-023-31341-0. Experimental study with 767 participants cited for the evidence on the shift in the user's moral judgment and the systematic underestimation of that shift.
Vallor, S. (2015). Moral Deskilling and Upskilling in a New Machine Age. Reflections on the Ambiguous Future of Character. Philosophy & Technology 28(1), 107–124. DOI 10.1007/s13347-014-0206-9. Origin of the concept of moral deskilling, cited in the block on the atrophy of the moral muscle.
Coeckelbergh, M. (2020). AI Ethics. MIT Press. Philosophical framework on the impossibility of extinguishing moral responsibility through delegation to systems, cited in closing the piece on the decision to delegate as a moral decision.
Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. Reference for the communicative strategy of the large commercial providers aimed at presenting their systems as neutral infrastructure.
Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. Can Language Models Be Too Big?. FAccT 2021. DOI 10.1145/3442188.3445922. Cited for the critique of the use of cognitive and agentive vocabulary applied to systems that have no real agency.
También te interesa
- Progressive trust
- Amplified confirmation bias
- Decision-making without data
- The emotional relationship with AI

Comments0
No comments yet.
Leave a comment