The black box as the norm. Electricity obeys physical laws; AI obeys whoever trained it

In this article

  1. The black box as a category of use
  2. The difference with generative AI
  3. Natural opacity, political opacity
  4. Protesting without an alternative
  5. What interpretability tries to scratch at
  6. The political consequence

Definitions · References · También te interesa · En otros sitios

I accept without protest that I don't know how it works, and I know why I accept it: protesting requires an alternative and I don't have one to hand. I'm told that AI's black box belongs to the same family as electricity or the combustion engine, things I use daily without understanding them. The comparison is rigged. Electricity obeys universal physical laws, public, reproducible by anyone with sufficient training. AI obeys whoever trained it, by private criteria, revisable at the operator's discretion, with no public access to the corpus, the method or the weights. The two opacities look alike from the outside —I don't understand what I use. Inside they're nothing alike.

The black box as a category of use

Before generative AI, the average adult in a developed country already lived alongside a handful of well-established black boxes. Electricity, the combustion engine, the antibiotic, the internet connection. Nobody loses sleep over not knowing what goes on inside a socket.

That calm was not pure ignorance. It rested on an implicit structure we rarely put into words: I trust the box because there are others who do understand it, because there's regulation that audits it, because there are public procedures for settling things when something goes wrong, and because I can fall back on an alternative if this particular one fails me. Four legs. Household electricity has all four: technical standards that regulate it, certified installers who audit it, a civil-liability regime that apportions blame and a market of providers under a common rule. The combustion engine too: type approval to sell it, a periodic inspection that checks it, a manufacturer that answers for its defects and dozens of models competing. The antibiotic passes through medicines agencies that approve it, pharmacovigilance that follows it afterward and a range of substitute therapies for almost any condition. The internet was built on open protocols and forces interoperability, with multiple providers and a partial net-neutrality regime in Europe.

The four boxes were black for the individual and transparent for the system that polices them. The opacity of use was admissible precisely because institutional opacity did not exist.

The difference with generative AI

Generative AI breaks those legs one by one, and not cosmetically.

Regulation arrives late and by halves. The European AI Regulation (EU 2024/1689) establishes the first serious framework, but its timetable lags behind deployment: the obligations for general-purpose models began to apply on 2 August 2025, general application is set for August 2026 and the toughest requirements, those for the high-risk systems of Annex I, don't enter until August 2027. By then, these tools will have been installed for years in the work of half the planet. External auditing, the second leg, simply has no equivalent. There's no inspection station for a generative model. The national AI safety institutes that have appeared since 2023 trial evaluations, but without real enforcement power. Nor are there public procedures for resolving a dispute over what the model produces: when an answer defames, plagiarizes or lies, the litigation ends up in ordinary courts armed with legal frameworks written with something else in mind.

That leaves the fourth leg, the alternative. And here the problem is the size of the pen. The general-model market is in the hands of a handful of companies, with some Asian competitor gaining weight. I can choose among them, sure, just as I can choose the colour of the car.

The four conditions that sustained the acceptance of opacity are eroded all at once. AI's black box has not earned trust by the same mechanisms that earned it for the earlier black boxes. We accept it in fact. Not by right.

Natural opacity, political opacity

The distinction that matters here isn't mine. The researcher Jenna Burrell formalized it in 2016, in an article on how the machine «thinks» published in Big Data & Society, where she separates three classes of opacity in machine-learning systems.

The first is opacity by secrecy, corporate or state: the owner of the system decides not to show how it works. It's political, it depends on a decision by someone with a name and a surname, and so it can be reversed with social pressure or with law. The second is opacity by technical illiteracy: the components are public, but most of the public lacks the training to read them. It's educational, and it's corrected by teaching. The third is the intrinsic opacity of the methods themselves: even with the code and the weights in front of you, the model's complexity overflows what a human brain can follow. It's structural, and it isn't fixed by revealing anything, because it's not a matter of concealment but of scale.

In today's commercial models all three are stacked. The user doesn't see the training corpus or the full weights; most wouldn't know how to read them if they had them; and the researchers who do have access decipher only a part of what the model does.

Electricity and the engine suffer only the second. The average user doesn't master electromagnetic induction or thermodynamics, but that knowledge is public and anyone with the right training can access, replicate and audit it. Generative AI carries all three at once, and it's the first —opacity by secrecy— that qualitatively separates it from the old boxes. That one is political. That one has an owner.

Protesting without an alternative

Protesting for real, not just talk, requires being able to leave. If system A strikes me as unacceptable and there's a system B that isn't, I switch and send a market signal. Without B, my protest is decorative.

With electricity, the car, the generic drug or the internet provider, that B exists, imperfect but working. With generative AI it narrows on two sides. Among providers there are differences of margin and focus, but they share the underpinnings: largely overlapping corpora, the same reinforcement tuning from human preferences, the same optimization measured by similar methods. Changing brand isn't changing architecture. And as for abstaining altogether, the door closes itself: in the professions that adopt these tools wholesale, whoever refuses is left at a disadvantage, and economic pressure ends up deciding for them.

Hence the uncomfortable consequence. Whoever is dissatisfied with the opacity has no protest mechanism with teeth. And their silence, the absence of visible protest, reads from outside as a yes.

It's worth not swallowing that reading. Acceptance is not consent. Acceptance is operational resignation in the face of a lack of exits. Consent is informed choice among real options. Democratic politics is built on the second; the market often runs on the first. Generative AI, today, operates with acceptance dressed up as consent, and the difference isn't semantic: it's the distance between being asked and being taken as asked.

What interpretability tries to scratch at

Against intrinsic opacity a whole field works, that of mechanistic interpretability, which aims to make auditable, even in pieces, what the model does inside. The advances are real and modest, both at once.

Chris Olah's team has been publishing since 2021 in Anthropic's Transformer Circuits Thread attempts to isolate the internal circuits responsible for specific behaviours. In October 2023, Bricken, Templeton and others released Towards Monosemanticity, where they applied sparse autoencoders —a technique that decomposes internal activations into more legible pieces— to a single-layer transformer, a toy model meant to test the method. The leap to a real model came in May 2024 with Scaling Monosemanticity, by Templeton and collaborators, which applied the same idea to Claude 3 Sonnet and reported extracting millions of identifiable features, some with demonstrated causal control. The example that circulated was the «Golden Gate Bridge» feature: forcing its activation, the model started talking as if it were the bridge. For its part, OpenAI's Superalignment team, dissolved and reformulated in 2024, had published in late 2023 work on supervising very capable models with weaker human signals.

All this qualifies intrinsic opacity. It doesn't remove it. The interpretability of a large model is today where neuroanatomy was in the seventies: circuits are located, a function is attributed to them, and a huge fraction of the system remains unmapped. There's a detail worth keeping in mind before getting carried away: these analyses don't touch opacity by secrecy. Access to the interpretation tools themselves remains selective, and the company that owns the model controls it. You can light up the box from inside and keep the door shut from outside.

The political consequence

Here I stop describing and give an opinion, though the opinion coincides with what the literature on algorithmic democracy has been formulating since before these models existed. Accepting AI as a black box comes dangerously close to accepting politics without democracy. The system makes decisions that shape public conversation, the labour market and the information that reaches us, without a control procedure comparable to the one we exert, more or less, over the State.

Frank Pasquale had anticipated it in The Black Box Society (2015), then regarding financial and reputational scoring algorithms, and the critique slotted effortlessly into generative models. His proposal was not total transparency, which would erode commercial incentives to the point of making the business unviable, but a qualified transparency: restricted but auditable access by regulators and by certified external experts. The idea is sensible and is partly captured in the European Regulation, which requires detailed technical documentation and systemic-risk assessments for general-purpose models, with access for the European AI Office. Whether that audit bites or only barks is something we'll see in the coming years, not before.

Electricity deserves the social deference we give it. So does the engine. So does the antibiotic. Generative AI will deserve it to the exact degree that its conditions of legitimation are built —regulation that arrives in time, auditing with teeth, real avenues for dispute, alternatives that aren't the same model with another logo. In the meantime, the only honest thing is to name the difference between the opacity of a socket and the opacity of a model, and not let the habit of not understanding the first make us sign the second in blank.

Definitions

Black box: a system whose internal workings are unknown to the observer, who can see only inputs and outputs.

Opacity by secrecy: opacity maintained by the owner's political or commercial decision, reversible through external pressure or regulation.

Intrinsic opacity: opacity inherent to the system's complexity, not removable by revelation because it exceeds human processing capacity.

Mechanistic interpretability: a field of research devoted to understanding which internal circuits of a learning model are responsible for which behaviours.

Sparse autoencoder (SAE): a technique that decomposes a model's internal activations into combinations of more interpretable features.

References

Burrell, J.How the machine «thinks»: Understanding opacity in machine learning algorithms. Big Data & Society 3.1 (2016). Source of the distinction between the three types of opacity.

Pasquale, F.The Black Box Society: The Secret Algorithms That Control Money and Information (Harvard University Press, 2015). Origin of the notion of qualified transparency.

Bricken, T. et al.Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread (Anthropic, October 2023). Application of sparse autoencoders to a one-layer transformer. https://transformer-circuits.pub/2023/monosemantic-features

Templeton, A. et al.Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread (Anthropic, May 2024). Application to Claude 3 Sonnet and the «Golden Gate Bridge» feature example.

Burns, C. et al.Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. arXiv:2312.09390 (OpenAI, 2023). Supervision of very capable models with weaker human signals.

Regulation (EU) 2024/1689European Artificial Intelligence Regulation. Application timetable: obligations for general-purpose models from 2 August 2025; general application in August 2026; high-risk Annex I systems in August 2027. https://artificialintelligenceact.eu/

Crawford, K.Atlas of AI (Yale University Press, 2021). Context on the political economy of AI systems.

También te interesa

En otros sitios

Comments0

No comments yet.

Leave a comment