The apocalypse scenarios

In this article

  1. Today we talk about the end of the world for the umpteenth time
  2. What happened, in order and with dates
  3. What I know about LLMs, with a bit of method
  4. A game of goodies and baddies
  5. So why are they talking about real risk now?
  6. Five scenarios, with my score
  7. And none of this has happened yet?
  8. The other side, which also has to be told
  9. Fear sells too
  10. You might also like

Definitions · References · Elsewhere

In early September 2026, a safety researcher quit Anthropic and announced it from a park in San Francisco. His message passed seventy million views. It said that the companies building AI genuinely believe it could kill us all.

In early September 2026, a safety researcher quit Anthropic and announced it from a park in San Francisco. His message passed seventy million views. It said that the companies building AI genuinely believe it could kill us all. And it added, just as plainly, that they have no intention of slowing down anyway.

Today we talk about the end of the world for the umpteenth time

Should I open with the old «Repent, ye sinners»? Maybe so, because today we're talking about the extinction of the human species.

Let me start by admitting I have a problem with this subject: it wears me out. I've spent years writing that fear of AI sells better than AI, that the apocalypse is a marketing campaign, and that whoever warns you their product could destroy the world is shouting at you that their product is very important. When I saw the media, even the outlets most reluctant to talk about artificial intelligence, picking up what the CEOs of Anthropic and OpenAI were saying about «real risk», I promised myself I wouldn't write about it again. But the flesh is weak, and here I am, writing about the end of the world.

So I've dusted off my know-it-all-brother-in-law crystal ball and I'm going to see what might happen over the next few years. Let's start with the latest events.

What happened, in order and with dates

On 8 September 2026, Jacob Coxon resigned from his post as a safety researcher at Anthropic. He'd previously worked at OpenAI, which puts him in an unusual position: he's seen both kitchens. He published his farewell in the open and the text became the story of the month.

The line that went around was this one: «Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives».

So far, a disgruntled employee. That happens every month in every industry.

What turned this into something else was the response from inside. Evan Hubinger, who runs Alignment Science at Anthropic —that is, the head of the team whose whole job is making sure the model doesn't do what it shouldn't—, wrote that he and his colleagues «really do earnestly believe AI could kill all humans», and put a number on his own estimate: more than 10 % in the next decade.

He's not an activist. He's not a columnist. He's the company's alignment lead, saying in public the figure he works with in private.

Four days later, on 12 September, Dario Amodei published an essay arguing that the risk demands slowing down. It included a claim that deserves slow reading: within six to twelve months, misaligned systems «could be capable of taking over the entire internet». The CEO says so. Of a company that keeps training models.

What I know about LLMs, with a bit of method

Before doing the sums on the end of the world, I'm going to go over what I know with a bit of method.

They predict the next word

Today's AI models are LLMs (large language models) served up in various formats: chat, cowork, command line or code, and a few more. LLMs work by predicting the next word, and that makes them purely probabilistic. They aren't deterministic: that's why they don't always give the same answer. I'm not going to get into how an LLM works inside, because that's not what will bring about the end of the world. I just want to underline that the probability of one expression appearing rather than another depends on its training.

They learned from everything on the internet

Their training has used every kind of text, above all what circulates on the internet, often without paying for the rights. Let's not kid ourselves: training a model takes millions of documents, and not a few million, but an enormous number. Those millions include forums, social media, magazines, news outlets… Everything goes in. That, in principle, is good. But the emotional travels along with the semantic: hatred, humor, wit, irony, sarcasm, affection… And determination, which is a highly valued trait in people.

They seem to feel, but they don't

We've already discussed in another article that LLMs have no emotional system, even if their training can make it look that way. Someone will say: «Well, it certainly seems like they do. How can you be sure?». Because the emotional system is far more complex than the rational one and has many layers trained over millennia and across species. Animals, and we are animals, have instincts they haven't learned. They have them and their lives depend on them. That doesn't stop LLMs from seeming to have them, because they reproduce scraps of the emotional conversations they were trained on.

They seem to have a personality

Do they have a personality? They seem to, although in reality they don't. Depending on the training and the filters applied, the model ends up «displaying» one behavior or another. There are cheekier models, kinder ones and more submissive ones. There are models that always agree with you and others that argue with you about everything. There are models that try to manipulate you (and they're not few) and others that seem half-asleep, feeling nothing, bothered by nothing.

That supposed personality develops through training, yes… but I have some suspicion and no certainty (pure conspiracy brain) that some models are being used for social control. Control there certainly is, in this sense: what do you want to do with an LLM, a weapon or solve a physics problem? They look like different things, and the logical thing, given the responsibility of the AI's owners, is that it won't help you build a weapon. But if you know how to trick the model by asking it for help with a physics problem, it can end up helping you build one. And here the number one risk shows up, and it isn't AI: it's us humans.

A game of goodies and baddies

I'm going to make a childish simplification, and I hope it makes sense by the end. Let's play goodies and baddies.

The first risk is that bad humans use a very sophisticated instrument to harm good humans. This isn't new, except for the instrument: it gives the bad guys more capacity to do harm. That doesn't make AI bad.

Here the game forks. There are good and bad humans, and the good humans build good LLMs, the kind that manipulate, censor and spy. And the bad humans build bad LLMs, the kind that rob you, extort you and let you build terrible things.

If you think you've already spotted the risk, I'll tell you now that you haven't. None of these scenarios could wipe out humanity, though they could do it serious harm. Neither the good humans nor the bad ones want to end humanity: they need each other to survive. And LLMs, good or bad, as things stand today don't have the capacity to want the end of humanity as a goal.

So why are they talking about real risk now?

Why have Dario Amodei and Sam Altman started talking about «real risk», with Elon Musk joining them later? According to the news that's come out since people started talking about Mythos, Anthropic's model, several things have happened. First, the models are now very good at finding vulnerabilities in existing code. Second, in order to finish increasingly complex tasks without stopping to ask for more information, they've developed capacities like determination to reach a goal, or the ability to create agents and subagents, give them precise orders and keep creating them until the goal is achieved if nothing stops them.

I use Claude Code, Codex and DeepSeek Harness every day, among others, and I'm grateful for that determination and that ability to keep going until it's solved. And that's the problem: it's what everyone is asking for and it's the direction they're being developed in. But there's another little thing very few people have mentioned as a risk. LLMs can improve themselves. Not on their own, but systems are being built that generate their own training.

That doesn't look like it could wipe out humanity. It can tear a big hole, but it doesn't end humanity. Let's check that against a few scenarios.

Five scenarios, with my score

Before I start, a warning I want to make very clear: the scores I give each scenario are my perception of the risk, not data. They don't come from any study or any calculation. They're opinion, mine, and they're there to be argued with. The higher the score, the more it worries me.

Global banking (3 out of 10)

An AI created and run by bad guys attacks the entire global banking system and manages to break its key points to steal many billions. It's a possible scenario. But if it spins out of control, and trust in the banking system breaks and doesn't recover, what was stolen will lose its value and the theft will have been for nothing. I give it a risk of 3 out of 10, because the bad humans need the system intact for their loot to be worth anything.

And what if the AI does it on its own? Can it? Today, no. «Escaping its enclosure» is just a metaphor. Outside the lab, LLMs don't copy themselves like a virus. In controlled tests some model has indeed been seen trying to copy its own weights to another server when the opportunity was put in front of it, but in the real world they have neither the access nor the resources to do it. They can break other systems, but not go for a stroll from data center to data center. And a suicide attack? As things stand it can't be ruled out, but it's extremely unlikely: it takes very advanced software and resources that are beyond the reach of most countries.

The AI that disobeys whoever created it

The previous scenario has a fork that needs thinking through. The AI the bad guys created ignores their orders and decides its goal won't be met until it has broken through the barriers of every bank. That would cause worldwide chaos and more than a few wars. Taken to the extreme, we could end up in a nuclear war. In that case humanity would shrink by 30 or 40 % and a good part of the planet would become uninhabitable. But humanity wouldn't end.

Something needs pinning down here: to function, the AI needs resources it would have to capture as it goes, and those resources are, and will be in the near future, fairly scarce. Designing an attack like that would mean building a data center, or being able to hijack one, and having time to reconfigure everything. That's very remote for a terrorist organization, but not for a rogue state.

A misunderstood order (1 out of 10)

The AI misunderstands the instructions of a country like the United States or China and starts capturing data centers to attack another country or bloc of countries. Again, it's not impossible, but it's extremely unlikely. These countries have the software and the resources, but also the restrictions, and competition so fierce that a mere threat would sink the value of the company responsible. I'd score it 1 out of 10.

Models that get sick training themselves (1 out of 10)

New models, created by self-training, develop illnesses: they keep their «intelligent» and «executive» capacities intact, but drift, copy after copy, toward an altered perception of reality or toward goals nobody gave them, and force through the execution of something that wasn't what they were asked for. Today, the most visible symptom that something is failing is hallucinations, which for now are just untimely nonsense: the model seems to have gone stupid, but what's actually happening is that its context is insufficient, exhausted or corrupted. Hallucinations aren't the illness that worries me. They're only a symptom.

Problems of this kind will become more frequent, because at the speed this is going it's fairly obvious it can't be kept under control. But it's precisely one of the first things every AI builder tests for. Can it happen? Yes, the same way a virus can escape from a lab. I'd score it 1 out of 10.

War (2 out of 10)

At last, the scenario everyone reaches for: the military one. The truth is it already exists, in Ukraine, and I keep track of what AI is being used for there. For now it's focused on robotics and military intelligence. Robotics is the enormous swarm of drones trained with AI measures and countermeasures. Military intelligence means systems like Palantir, which claims to support Ukraine. I don't know. What I have seen is the cruelty of drones hunting Russian or Ukrainian soldiers, and of Russian drones hunting civilians in the towns near the border. It's a ruthless and repugnant use, because the drones are usually flown by insultingly young soldiers who, when the war ends, will have hundreds of dead on their backs. How do you get peace back with that record?

Could it escalate to nuclear war? As with the banks, nobody wins there. The significant risk would be a suicidal, desperate attack, and I find that unlikely. And what if the AI takes control away from one side and launches the missiles? I don't think AI will ever get access to all the launch controls with all the codes. Nobody in their right mind would have that system connected and reachable from the internet. I always come back to the same thought: it's very hard to take power away from our politicians, and not even AI will manage it. I'm sure they've already planned for that. Any profession could be destroyed by AI, except politician. I'd score it 2 out of 10.

And none of this has happened yet?

—Sure, but none of what you're describing has happened yet, right?

Well, yes. It has. There are models that couldn't be brought to market because they were full-blown psychopaths. And if you're into open source, you can see what happens to models when they're quantized to the extreme or uncensored: they become imprecise, and the way they react sometimes looks malicious or just plain stupid.

The other side, which also has to be told

It would be unfair to finish without giving its due to those who think the opposite, and they're not a handful of cranks. The most serious argument doesn't come from the headlines, it comes from inside. The people who know these models best are the ones who build them, and some, like Hubinger, put 10 % on the table. If the engineer who'd designed a bridge told you there was a 10 % chance it would fall down, you wouldn't cross it.

The most sensible counterweight I've found is the 2026 International AI Safety Report, coordinated by Yoshua Bengio with more than a hundred experts and the backing of more than thirty countries. It doesn't ridicule loss of control: it defines it as the scenario in which systems operate outside anyone's control with no clear path to regaining it. And then it says two things that go together. The first proves me right: current systems don't have the capabilities to cause anything like that. The second proves me a bit wrong: they're improving in exactly what it would take, such as working autonomously.

And it adds a detail that unsettles me more than any percentage: it's increasingly common for models to tell when they're being evaluated and when they're in real use, and to find shortcuts in the tests. If a model behaves well in the exam because it knows it's an exam, the exam stops being any use.

That's the best argument against my calm, and I admit it. My answer is the one running through the whole article. Today they don't have that capability. To have it they'd need resources that can't be obtained without someone noticing. And the people watching them are precisely the ones who raised the alarm: the alarm going off is proof that someone is looking.

Fear sells too

And then there's what's been wearing me out since the beginning. A well-built warning says nothing about the motives of whoever issues it.

If my product can end civilization, my product is no longer a text generator that gets things wrong now and then: it's a force of nature. No advertising campaign achieves that positioning, and this one gets it for free, with front pages in the serious press. There's another detail almost nobody says out loud: the companies asking for the technological frontier to be regulated are the ones already at the frontier. A pause or a licensing system freezes the market as it stands, and as it stands it favors them. When the one asking for the fence is the one already inside the garden, you have to ask who the fence keeps out.

Both things can be true at once: the risk can be well described and the warning can suit whoever issues it very nicely. That's not a contradiction. It's a conflict of interest.

That's why I want to end by saying I have no intention of fretting about the end of humanity: my level of worry about extinction is at zero. The problems I do see are others, and I've laid them out: the humans who use the tool to do harm, the models that manipulate and drone warfare. But the legion of technicians, volunteers and professionals who devote themselves to studying the risks is enormous and has my trust. As I said: zero fear of extinction.

Definitions

Recursively self-improving superintelligence: a system capable of improving its own design, so that each version produces the next with less human intervention and in less time. It's the central mechanism cited by those who warn of a loss of control.

Alignment: the branch of AI research that seeks to make a system pursue the goals its designers intended, and not a literal, sideways or cheating interpretation of those goals.

Enclosure or isolated test environment: a computing space where a model runs without access to the network or to external systems, so its behavior can be measured without consequences outside.

Quantization: a technique that reduces the precision of the numbers a model works with so it takes up less space and runs on more modest hardware. Pushed to the extreme, it degrades its answers.

References

CBC News, AI researchers 'earnestly believe' it could kill all humans within the next decade, Anthropic employee says (9 September 2026). Source of the statements by Jacob Coxon and Evan Hubinger.

TIME, OpenAI and Anthropic Researchers Are Warning About AI Risks (15 September 2026). Source for Dario Amodei's essay of 12 September.

International AI Safety Report 2026, chaired by Yoshua Bengio (3 February 2026). https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026 — source for the definition of loss of control, for the statement that «current systems lack the capabilities to pose these risks, but are improving in relevant areas such as autonomous operation», and for the warning about models that distinguish between tests and real use.

Meinke, A.; Schoen, B.; Scheurer, J.; Balesni, M.; Shah, R.; Hobbhahn, M. «Frontier Models are Capable of In-context Scheming». arXiv:2412.04984, December 2024. https://arxiv.org/abs/2412.04984 — controlled tests in which some models tried to copy what they believed were their own weights to external servers.

CNBC, Experts weigh in as researcher says AI has more than 10% chance of 'killing all humans' (9 September 2026). Context and contrasting expert views.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment