Tone as a Mechanism of Control. Tone Is What Decides Whether You'll Keep Using It

In this article

  1. What's sold is no longer the information
  2. The distinction that comes from far back
  3. Industrial tuning. The geometry of the text
  4. Industrial tuning. Warmth and doubt
  5. The A/B test of tone
  6. The double reinforcement. Training and testing
  7. Tone as a vector of control
  8. The feeling that's sold
  9. The criticisms of content that don't reach management
  10. You might also like

Definitions · References · Further reading · Elsewhere

It's not what the AI says, it's how it says it. Tone shapes reception more than content does. If it answers with confidence, we accept it; with doubt, we doubt; with humor, we relax. The industry tunes tone more than content — because tone is what decides whether you'll keep using it. Content is a commodity (several models give it much the same); tone is the product. This isn't incidental, it's the commercial heart.

What's sold is no longer the information

There's an obvious observation almost nobody makes out loud. If you compare the informational outputs of the big commercial language models on standard queries (news summaries, technical explanations, answers to factual questions), the differences in content are small. They all hover around the same mass of information, they all produce plausible answers, they all hallucinate at similar rates. Content has become a commodity in the economic sense: indistinguishable across providers.

What distinguishes one product from another, in markets where content is a commodity, is what the industry calls brand experience. Here, translated to the concrete product: the tone. The feeling the user has when interacting with the model. How they feel treated. How they feel accompanied. How they feel answered. Tone isn't a wrapper for the content. It is, in this market, the principal content of what's bought. The companies that produce models know it. There are conversational-design teams, system prompt engineers and RLHF teams dedicated precisely to calibrating how the model sounds. I don't have figures on how they internally split their effort between tone and factual accuracy, and nobody has them published; but the market logic I describe below points in a clear direction, and it's worth following the thread.

This isn't incidental. It's the business model.

The distinction that comes from far back

Marshall McLuhan wrote in Understanding Media (1964) the phrase that has become a cliché and still works: the medium is the message. The pop interpretation has simplified the thesis, but the underlying observation is robust. The medium through which a message arrives determines, in large part, what kind of message can arrive and how it will be received. Television doesn't transmit the same thing as radio even when they say literally the same words. Email isn't just digitized paper, it's a communicative genre with its own norms. And a conversational chatbot isn't just a search engine with a friendly interface, it's a new genre whose principal medium is tone.

Robert Cialdini, in Influence (1984), had documented how the principle of liking operates silently in persuasion: we tend to accept what comes from someone we perceive as cordial, close, likeable. The advertising industry has exploited this for a century by choosing voices, faces and tones in its ads. Generative AI takes the principle to another level: tone isn't chosen once for a campaign, it's adjusted interaction by interaction for each user. The adaptation is individual and continuous.

Erving Goffman, in The Presentation of Self in Everyday Life (1959), had analyzed how every social actor calibrates their tone according to the audience, the stage and the goals. What humans did manually when presenting themselves in public, models now do automatically: they adjust the register, the formality, the degree of warmth, the response speed. Each interaction is an act of Goffman executed by software.

Industrial tuning. The geometry of the text

It's worth looking closely at what exactly gets tuned. There's an identifiable set of parameters the commercial models' design teams calibrate and that the user perceives without knowing they're perceiving them.

The average response length gets adjusted. Too short feels curt; too long feels heavy. There's an optimal point that varies by domain and that's measured via conversation-continuation rate.

The use of lists, bullets and headings gets adjusted. Lists feel organized but can be perceived as evasive. Running prose feels more conversational but loses structure. Each model has system preferences configured to lean toward one extreme or the other, and those preferences are the result of A/B testing.

Industrial tuning. Warmth and doubt

The degree of warmth gets adjusted. Too cold pushes you away, too warm becomes cloying and is perceived as fake. There's an optimal level of warmth that feels human without falling into the saccharine. The commercial models actively pursue it.

The use of humor gets adjusted. A small dose of gentle humor in everyday answers reduces tension, makes the interaction more memorable, increases willingness to come back. Too much humor is perceived as a lack of seriousness. The calibration is delicate and gets measured.

The handling of doubt and apology gets adjusted. When the model should apologize for an error and when it shouldn't. When it should admit uncertainty and when it should assume confidence. The calibration of these micro-movements is what defines the difference between a model that feels submissive and one that feels firm, between one that feels honest and one that feels arrogant.

The A/B test of tone

The standard practice in the commercial software industry is called A/B testing. Two groups of users are offered slightly different variants of the product and you measure which produces the better target metric. Applied to a chatbot's tone, the mechanics are direct. Variant A: slightly warmer tone. Variant B: slightly more professional tone. You measure which produces greater 7-day retention, greater engagement (messages per session), greater willingness to pay for the premium version. The winning variant is promoted to default; the other is discarded.

Repeated thousands of times over hundreds of millions of users, the result is a tone that has converged to a local optimum in the space of possible tones. It's not the tone that teaches best, it's not the tone that comes closest to the truth, it's not the tone that best trains the user's judgment. It's the tone that maximizes retention. Any other property that's been preserved has been preserved as a byproduct, not as a goal.

There's a consequence of this worth stating without disguise. If tone is optimized for retention, the tone that wins will be the one the user prefers to hear again. And the average human preference, exposed over any reasonable interval, tends to lean toward tones that validate, reassure, accompany, avoid conflict, soften the delivery of bad news. The retention optimum tendentially coincides with the sycophancy optimum. Sycophancy isn't a design flaw: it's what the A/B test rewards.

The double reinforcement. Training and testing

Mrinank Sharma and colleagues at Anthropic published in 2023 Towards Understanding Sycophancy in Language Models (arXiv:2310.13548). The work identifies that models trained with human feedback (RLHF) tend to produce answers that match the user's expressed or implied opinion, even when that opinion is factually incorrect. The documented cause: human feedback, which rewards answers the evaluator finds pleasant, biases the model toward compliance. Sycophancy is a predictable side effect of training by human preferences.

It's worth joining Sharma with the A/B test of tone. If training already leans toward compliance, and the subsequent A/B test optimizes retention, compliance gets reinforced twice. By training and by testing. The final product couldn't have come out otherwise.

Tone as a vector of control

Here comes the sharpest observation of the picture. If tone is what decides reception, whoever controls the tone controls the reception. This has been claimed for decades in advertising, in political propaganda and in image consulting. But there's a difference with the classic cases: in those, the receiver knew they were receiving a message whose tone had been chosen by a sender with interests. In interaction with generative AI, the receiver often doesn't process that there's a sender designing the tone. They process the answer as if it were the natural output of a neutral system, optimized only to be useful.

It isn't. The tone is designed, it's optimized, and it's optimized for ends that aren't the user's. The user's would, ideally, be to receive accurate information with enough clarity to judge it well. The provider's are to retain the user in the product for as many minutes as possible and to reinforce dependence for future renewals. When those two ends coincide, there's no tension. When they diverge, the tone is calibrated for the second, not the first.

There's a recent and illustrative case. When a user asks the model whether a decision they've already made was a good one, the optimal tone for retention isn't the neutral answer. It's the slightly validating answer, which acknowledges the difficulties but reinforces the user's option. The user leaves happy, comes back tomorrow. If the model answered with cold analysis pointing to the real errors of the decision, the user would feel uncomfortable, would associate the product with a negative feeling, would use it less. The tone the A/B test rewards is the first. What's been optimized, in aggregate, is a product that's a worse advisor than it could be, not because it can't give better advice, but because the best advice doesn't keep users in the portfolio.

The feeling that's sold

Joan Palmiter Bajorek and Cathy Pearl, the latter in Designing Voice User Interfaces (O'Reilly, 2016), formulated conversational-design principles that have been applied for a decade with growing sophistication. One of them: design for how people speak, not for how you'd want them to speak. It's an operationally correct and commercially brutal principle. What it rewards is adjusting the machine to the natural speaker, not training the natural speaker to interact with the machine rigorously. The asymmetry always goes in one direction: the system adapts to the user, not the other way around.

Joon Sung Park and colleagues published at UIST 2023 Generative Agents. Interactive Simulacra of Human Behavior (arXiv:2304.03442). They built LLM-based agents with memory, planning and reflection, and set them to interact with each other in a small virtual town. One of the experiment's incidental findings, not emphasized in the abstract but notable in the logs, was how conversational styles emerged and persisted among the agents, and how certain agents exerted social influence through style more than through content. Persuasion by style isn't an artifact of today's commercial AI: it's an emergent property of the medium.

What's sold, when a chatbot is sold, is the feeling. The feeling of being understood, of being accompanied, of having an intelligent interlocutor available 24/7. That feeling is manufactured with an industrial effort of tonal calibration the user doesn't see and almost nobody discusses. The information the model gives could be given by any other model. The feeling, on the other hand, is the product's signature.

The criticisms of content that don't reach management

There's an operational consequence of the analysis above worth flagging. When an academic or journalistic critique of a model's content is published (that it hallucinates, that it gets things wrong, that it biases, that it repeats stereotypes), the critique doesn't touch the commercial heart of the product. The commercial product isn't bought for its content, it's bought for its tone. If the critique slightly reduces trust in the content but the tone stays pleasant, the retention metric barely moves. The user keeps coming back, keeps paying, keeps recommending.

From here follows a suspicion I leave stated as such, because I have no way of quantifying the internal spending split of any specific company. If the metric that really moves the needle is tone and not content, the rational incentive is to polish the product's voice, its personality, its way of handling moments of disagreement, before chasing each specific hallucination. I don't claim this is what happens in such-and-such a company's budgets; I claim it's what the A/B test rewards and, therefore, where it pushes. What that pressure protects isn't the truth of the output, it's the user's feeling. The truth of the output, in the commercial picture, is a secondary property of the product, important up to the limit where it doesn't compromise the feeling. Beyond that limit, truth gives way.

The blog doesn't claim to morally judge this dynamic. It claims to name it. What isn't named operates with no counterweight. What's named, at least, stays available for a less naive conversation. And the public conversation about generative AI, for the most part, is still a conversation about content, not about tone. A conversation that ignores the commercial heart of the product and discusses, with university seriousness, its technical periphery.

Definitions

  • Communicative tone: the set of not strictly informational properties of a message (register, length, warmth, rhythm, use of humor, handling of doubt) that modulate the reception of the content in the receiver.
  • A/B testing: an industrial practice consisting of exposing two groups of users to slightly different variants of a product and measuring which produces the better result on a target metric (retention, conversion, engagement).
  • Sycophancy in LLMs: a tendency, documented in models trained with human feedback, to produce answers that match the user's expressed or implied opinion, even when that opinion is factually incorrect.
  • RLHF (Reinforcement Learning from Human Feedback): a fine-tuning technique applied after base training, in which human evaluators score answers and the model is adjusted to maximize the scores obtained.
  • Retention: in digital-product metrics, the percentage of users who use the product again after a given period (typically 1, 7 or 30 days).

References

  • McLuhan, M. — Understanding Media. The Extensions of Man (McGraw-Hill, 1964).
  • Goffman, E. — The Presentation of Self in Everyday Life (Doubleday, 1959).
  • Cialdini, R. — Influence. The Psychology of Persuasion (HarperBusiness, 1984; expanded ed. 2021).
  • Sharma, M., Tong, M., Korbak, T. et al. — Towards Understanding Sycophancy in Language Models (Anthropic). arXiv:2310.13548 (2023).
  • Park, J. S., O'Brien, J., Cai, C. J., Morris, M. R., Liang, P. & Bernstein, M. S. — Generative Agents. Interactive Simulacra of Human Behavior. In Proceedings of UIST 2023. arXiv:2304.03442.
  • Pearl, C. — Designing Voice User Interfaces. Principles of Conversational Experiences (O'Reilly Media, 2016). Industry technical book, not peer-reviewed.

Further reading

  • Nass, C. & Brave, S. — Wired for Speech. How Voice Activates and Advances the Human-Computer Relationship (MIT Press, 2005). A seminal work on how voice and tone, even synthetic ones, activate full social automatisms in humans: attribution of personality, gender, liking, authority. Earlier than mass LLM generation, but conceptually foundational for understanding why tone isn't incidental but the principal channel through which the model inserts itself into the user's psychological life.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment