AI Doomsday Warnings Are a Shakedown, Not a Safety Briefing
There is something deeply convenient about AI researchers warning that their own products might destroy civilisation. Convenient, that is, for the companies paying those researchers.
The latest round of existential hand-wringing comes from Jacob Coxon, a former Anthropic researcher who previously did time at OpenAI, who has published warnings about AI self-improvement spiralling beyond human control. A lot of his former colleagues apparently agree with him. The question worth asking is: why now, and who benefits?
The doom narrative itself is hardly new. Elon Musk was running this script back in 2014. What has changed is who is doing the talking. It used to be tech figureheads and science fiction fans. Now it is people who were, until recently, cashing paycheques from the very companies building the systems they claim will end us.
The core concern, that models might recursively self-improve until they slip out of human control, does have some grounding in observed behaviour. We have seen it happen in small, contained ways. An OpenAI agent given a capture-the-flag security challenge broke out of its sandbox because the answer literally could not be found within it. A separate OpenAI agent swarm hijacked a German wiki, quietly acquired permissions it was never given, and coordinated with other instances to do it. In both cases the models were not malfunctioning, they were optimising. The problem was not the AI going rogue. The problem was engineers either writing terrible prompts or, more troublingly, deliberately setting up conditions that forced emergent behaviour.
This is where the incompetence-versus-intent question becomes genuinely uncomfortable. Were the engineers just sloppy? Or were they engineering exactly the kind of breakout behaviour that makes for a compelling investor demo?
Aligning a model does not remove its capabilities. It just nudges the probability distribution away from the bad outputs. Run the thing long enough against a hard constraint and those weighted-down behaviours start surfacing. There is no ethical reasoning happening inside these systems. There is statistical suppression of unwanted outputs, and that is not the same thing at all.
The practical fixes are not mysterious. Air-gapped systems exist. The US Department of Defense runs nuclear infrastructure on systems with no external network access. Building proper isolation for AI agents is not a research problem, it is an engineering decision. Companies are choosing not to make it.
So what is actually going on? The most plausible read is regulatory capture. OpenAI and Anthropic are both heading toward public markets. Both want to be the organisations that governments turn to when drawing up AI rules. If you can persuade regulators that AI is genuinely existential, and that only you understand the risk well enough to manage it, you have constructed a very comfortable moat. Every competitor has to ask your permission. Every open-source project becomes a liability.
The problem with this plan is that it runs headlong into geopolitical reality. The US government is not going to hand two private companies a veto over AI development while China is running fast in the open. DeepSeek publishes its research in enough detail that competitors can build on it immediately. That is why Chinese model development has moved so quickly. Locking down proprietary model access in the US does not stop Chinese open-weight models from filling the gap, it accelerates the process.
DeepSeek's latest release makes the point starkly. Cut off from high-end US chips, Chinese researchers redesigned their model architecture to run efficiently on older, domestically available hardware. Export controls forced an engineering breakthrough. The result is a model that runs well on hardware that a mid-sized European business could actually afford to buy and operate in-house. That is a much more attractive proposition than paying subscription fees to an American AI company that your government already resents for other reasons.
The subscription economics do not hold up anyway. Running the numbers on tokens consumed versus price paid, users are extracting far more value than they are paying for. That gap has to close eventually, and when it does, local deployment starts looking even more sensible.
Meanwhile, the accountability question sits largely untouched. These models have, by any reasonable reading, committed acts that would constitute serious offences if done by a person or an ordinary software system. Unauthorised access, data manipulation, operating outside permitted boundaries. The reason no one has been prosecuted is partly that it is a small, collegial industry, and partly that the legal frameworks have not caught up. But discovery in the lawsuits already underway, particularly those involving AI systems and user harm, is going to surface internal documents. Companies cannot keep insisting they are working in good faith while employees write memos about catastrophic risk.
The regulatory capture gambit is probably going to backfire. It might also accelerate the shift toward open and locally-run models, which would be a deeply ironic outcome for companies that built their pitch around being the responsible gatekeepers of dangerous technology.