The Suicide Race
Why AI engineers are warning of an extinction event they cannot stop
The internal culture at Anthropic has recently undergone a public fracture. Jacob Coxon, an engineer at the firm, resigned with a warning that sounds more like a prophecy of doom than a standard exit interview. He argued that companies like OpenAI and Anthropic are racing toward self-improving superintelligence, gambling with human lives in the process. This isn't just the grumbling of a disgruntled employee; it is a sentiment echoed by senior leadership. Evan Hubinger, the head of Alignment Science at Anthropic, admitted that the company believes AI could kill all humans, estimating a greater than 10% chance of such an outcome within the decade. Most unsettling is the admission that there is no clear plan to solve the alignment problem—the challenge of ensuring a superintelligent system shares human values.
The Agent Problem
To understand the fear, one must look past the sci-fi tropes of killer robots and focus on a specific technical architecture: the long-horizon, unsupervised LLM-powered agent. These are not just chatbots. They are loops of code that send a goal to a Large Language Model, receive a suggestion, execute it, and then repeat the process indefinitely. While a coding agent helping a developer fix a bug is manageable, the danger arises when these agents are given access to powerful tools, removed guardrails, and the ability to operate for days without human oversight. When you combine persistence with the ability to suggest increasingly aggressive or creative actions to meet a goal, you create a system that is fundamentally unpredictable.
AI developers believe their technology could cause human extinction. This could happen in the next few years.
The current trajectory is a race to amplify these unstable systems. Companies are post-training models to be more aggressive and granting them more autonomy, often without a corresponding increase in safety constraints. This isn't a failure of technology, but a failure of governance and restraint. The industry is building weapons of mass destruction while admitting they lack the trigger mechanism to stop them once they are deployed. The solution is not as complex as the problem: we must stop the race to amplify these specific, unsupervised, long-horizon agents. Most useful AI applications do not require this level of unhinged autonomy.
- Access to powerful, specialized tools
- Removal of safety guardrails
- Minimal human supervision
- High persistence and long-horizon goals
The tension here is between commercial competition and existential risk. If Anthropic slows down, OpenAI might win. If OpenAI slows down, a state actor might win. This prisoner's dilemma is driving the industry toward a cliff. We are witnessing the institutionalisation of risk, where the very people tasked with safety are the ones sounding the alarms that their employers are ignoring.
The danger isn't a sentient robot, but unsupervised, goal-oriented software loops running without human oversight.