The illusion of control. The prompt is input, not an order

In this article

  1. The interface that suggests command and the mechanics that deny it
  2. Prompt engineering as an industry of opacity
  3. The informational asymmetry
  4. The system prompt, that contract you didn't sign
  5. The reranker that rewrites your prompt
  6. The black box that interpretability doesn't cure
  7. The operational question: what to do with this?
  8. Four stances toward opacity
  9. You might also like

Definitions · References · Elsewhere

You think you're steering the system because you write prompts and get answers back. But the system decides what it understands, what it ignores, what it prioritizes, what it rewrites internally, what it blocks for moderation, what it answers from the median of the corpus. Your prompt is input, not an order. The human-AI conversation has the appearance of human control and the substance of an opaque negotiation. This isn't fixed with better prompting —it's fixed by accepting there is no real control.

The interface that suggests command and the mechanics that deny it

When you type into the text field and hit send, the interface returns an answer. The operation has the phenomenology of an executed order: you give the instruction, the system obeys. Because that format is the same one your operating system uses when you tell it to open a file, you read it as a technical command. The reading is wrong.

Between the moment you hit send and the moment the answer appears, at least seven operations happen, in order, that you have no direct information about and no control over. It's worth listing them, because the list, by itself, makes the argument.

Your text gets tokenized: it's broken into sublexical units the model knows. The segmentation isn't yours; in the word "extraordinarily," the model might see "extra," "ordinari," "ly" or any other cut. What the model processes isn't your words, it's its tokens.

Your text gets prepended with an invisible system prompt. Every commercial product —ChatGPT, Claude, Gemini, Copilot— adds, before your text, a long instruction, configured by the company, that defines the tone, the restrictions, the prohibitions, the persona. That instruction is contractual to the product and you haven't seen it. In some cases jailbreaks have revealed fragments. The full version, normally, is known only to the maker.

Your text gets filtered by input moderation. If it contains patterns the product's policy flags as sensitive, it can be blocked before it reaches the model, or rewritten by an intermediate layer, or annotated with metadata that will change the later answer.

Your text may be rewritten by a reranker or reformulator. Products like DALL-E 3, for a while, explicitly rewrote prompts before passing them to the generator. Some RAG systems reformulate your query to optimize the internal search. What reaches the model, in these cases, isn't what you wrote.

The model decides what to do with the input. This decision isn't deterministic: there's temperature, top-k, top-p, sampling parameters the company configures. The same input can produce different answers at different times.

The answer gets filtered by output moderation. If the model produced something the product doesn't want to appear, it's swapped, reformulated or replaced with a generic message.

The answer gets post-processed: formatting is added, disclaimers are stripped, length is adjusted, style templates are applied.

Of these seven layers, the user controls only the first of the seven, and even then only partly (because the tokenization is the model's choice). The other six are the maker's decisions. Calling the operation "giving an order to the system" is generous. Marcus and Davis already pointed out in Rebooting AI (2019) that the operational opacity of these commercial systems —not being able to look inside, not being able to audit without internal access— turns any talk of user control into rhetorical courtesy.

Prompt engineering as an industry of opacity

There's an entire discipline, prompt engineering, that has emerged to mitigate this structural ignorance. Wei and others, in Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022, arXiv 2201.11903), published the foundational paper. They showed that adding instructions like "think step by step" to the prompt substantially improved the model's performance on reasoning tasks. The observation is real and replicable, and it opened a field of practice.

What's worth looking at slowly is the epistemological structure of that field. Prompt engineering isn't engineering in the strong sense. It doesn't design the system; it probes it. It discovers, by trial and error, which formulations produce the best results in the specific model the user has in front of them. It's the discipline of learning the reflexes of a black box the system doesn't document. Each company publishes partial guides —Anthropic, OpenAI, Google have their respective prompting cookbooks— but none documents the entirety of the system prompt, nor the moderation logic, nor the internal rerankers, nor the post-processing layers. The official guide is already, by construction, part of the product's marketing, not a complete technical description.

The ceiling of prompt engineering, in this sense, is the system's opacity. However well a user learns to probe, there are structural limits: silent changes to the system prompt between product versions, moderation adjustments updated without warning, rerankers being retrained, samplers changing configuration. The skill the user accumulates has an expiry date and nobody will warn them about it.

The informational asymmetry

There's a property of the human-AI conversation worth naming because it appears in no human-human conversation. You share context. The system doesn't.

When you write a prompt, you give the model: your situation, your apparent intention, your level of knowledge, your emotional state, the domain you're operating in. Sometimes you give it explicitly, sometimes it's inferred from the style. The model, meanwhile, gives you nothing equivalent. It doesn't tell you which version it is. It doesn't tell you its cut-off date. It doesn't tell you which moderation layers fired in your session. It doesn't tell you whether your prompt was rewritten before reaching it. It doesn't tell you how confident it is in the answer. It doesn't tell you what sources it drew on. It gives you the output and, optionally, a generic disclaimer along the lines of "I'm a language model and I can make mistakes."

This asymmetry is the signature of a negotiation. In a negotiation, the parties exchange information strategically. The party with more information extracts more value. In the human-AI conversation, one party has zero information about the other and the other has all the information the first chose to share. The asymmetry is structural, not accidental. And the negotiation, within that asymmetry, has a predictable outcome: the system sets the frame, the user operates within the frame.

Crawford, in Atlas of AI (Yale UP, 2021), framed the asymmetry in political terms. What the interface sells as a conversation between equals is a relationship between centralized corporate infrastructure and an individual user with no means to audit it. The verb "talk to the AI" papers over what is materially interacting with a service run by a company whose configuration changes without your informed consent.

The system prompt, that contract you didn't sign

It's worth pausing on the system prompt, because it's the cleanest case of invisible control. Every commercial product has one. The instruction the model receives before your message describes who it "is," how it should answer, what it can't say, in what tone, with what level of detail. That instruction was written by the company. You didn't write it. But the answer you receive is shaped by it as much as or more than by your own prompt.

When a jailbreak reveals a fragment of the system prompt —this has happened with several products through 2023-2025— the content usually includes: the persona the model must simulate, the topic areas to avoid, the response reflexes to certain triggers, the instructions for presenting uncertainty, the commercial directives ("don't recommend competitors' products"). The user who writes "be honest" is directing an instruction to the model whose priority relative to the system prompt is predetermined by the company. Your requested honesty competes against the nuances the company already decided.

There are more sophisticated versions. Some products change the system prompt according to the context inferred about the user: location, language, subscription tier, usage history. The negotiation becomes, in these cases, personalized without your participation. You get a version of the model configured for you according to criteria you don't control and that aren't explained to you. The asymmetry deepens.

The reranker that rewrites your prompt

There's a particularly revealing layer in some products: the prompt rewriter. DALL-E 3 made it explicit for a while. When the user asked for "a woman reading under a tree," the system internally rewrote the prompt into something like "a woman of about thirty with brown hair, sitting under an oak tree with an open book, golden afternoon light, realistic photographic style," and passed that expanded version to the generator. The intent was to improve output quality. The effect, from the user's side, is that what your model generated wasn't what you asked for. It was what the rewriter decided you'd have asked for if you'd been more specific.

The operation is defensible as a UX improvement. It's also, by construction, silent interpretation of your intention. If the rewriter decided that "woman" meant "a young, white, Western woman," that decision came out of the rewriter's weights, not your prompt. And the rewriter's weights carry the biases of the corpus it was trained on.

Commercial RAG systems do something similar. Your question gets reformulated before the search against the index. The reformulation optimizes the search but discards nuances of the original. What the model ends up answering is shaped by the reformulation, not by your question. And the reformulation isn't shown to you.

The black box that interpretability doesn't cure

There are serious research lines working on internal transparency. Anthropic published in 2023 and 2024 mechanistic-interpretability papers —Towards Monosemanticity, the work on circuits in small and mid-sized models— that decompose, up to a point, what the models do inside. The field is advancing. What advances, though, is the scientific understanding of the model, not the commercial product's transparency to the end user. Knowing technically that a more transparent model could be built isn't the same as having access to that transparency when you use the product.

The distinction matters. The commercial product will remain, for reasons of competition and intellectual property, a black box with an unpublished system prompt, undocumented moderation layers and unexplained internal rerankers. The interpretability researched in labs doesn't automatically translate into a control panel accessible to the user who decides to move from an open-source model to a commercial service. The opacity isn't a technical flaw. It's part of the business model.

The operational question: what to do with this?

There's a temptation to close the observation with a recipe. The temptation is honest and worth partly resisting, because the recipe being sold —"learn better prompting"— is precisely part of the problem. There's no way to have real control inside an architecture designed not to grant it. What there is are more or less lucid operational stances toward the situation.

The naive stance treats the prompt as an order and is surprised when it isn't carried out. It's the stance of the user who enters the product believing they have technical command and leaves frustrated when the system refuses something, or rewrites something, or ignores something.

Four stances toward opacity

The technical stance accepts the opacity and probes it methodically. It's the professional prompt engineer. It squeezes the maximum possible performance out of the margins the system grants. It knows those margins can change tomorrow without warning.

The political stance recognizes that the problem isn't one of individual command but of governance. It argues, as Bender and others do (FAccT 2021), that the situation demands transparency regulation: an obligation to publish the system prompt, the rerankers, the moderation layers, the personalization criteria. That regulation, today, doesn't exist in enforceable operational terms.

A fourth stance —probably the only one that protects the individual user from self-deception— is the stance of someone negotiating with an opaque interlocutor. Assuming there's no command, only interaction. That each answer is the result of a negotiation whose terms you don't control. That the language of "control," "order," "instruction" mistranslates what's happening, and that the correct language is "request," "suggestion," "input." Calling your input input —not an order— is the one small vaccine you have against the illusion the product sells as an interface.

Definitions

System prompt. A long instruction every commercial product prepends to the user's message before sending it to the model. It defines persona, restrictions, prohibitions and tone. Not usually published, accessible to the maker.

Tokenization. The operation of segmenting text into sublexical units the model recognizes. It doesn't match human words. The maker's decision, not controllable by the user.

Reranker / prompt rewriter. An intermediate layer that reformulates the user's input before passing it to the model. Frequent in commercial products to optimize output or search quality. It introduces silent interpretation of the user's intention.

Input and output moderation. Filters each product applies to what the user sends and what the model produces. They can block, rewrite or annotate. Criteria configured by the maker, not published in detail.

Sampling. The procedure by which the model chooses the next word among the options its probability distribution allows. Parameters like temperature, top-k and top-p configure the sampling and the product decides them, not the user.

Mechanistic interpretability. A research line that studies the internal workings of models by decomposing their circuits. It advances technical understanding of the model; it doesn't, on its own, produce transparency of the commercial product to the end user.

References

Anthropic (2023). Towards Monosemanticity. Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. A reference work in the mechanistic interpretability of language models.

Bender, E., Gebru, T., McMillan-Major, A. & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021. An explicit call for transparency regulation of commercial LLM systems.

Crawford, K. (2021). Atlas of AI. Power, Politics, and the Planetary Costs of Artificial Intelligence. Yale University Press. A political framework on the infrastructural asymmetry between the user and commercial AI providers.

Marcus, G. & Davis, E. (2019). Rebooting AI. Building Artificial Intelligence We Can Trust. Pantheon. An early warning on the operational opacity of commercial systems and the difficulty of auditing them without internal access.

Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. arXiv: 2201.11903. The foundational paper of prompt engineering as a practice of probing the model.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment