After Coxon’s AI extinction warning, Habryka says the ‘how’ has already been mapped
New Delhi: Days after former OpenAI and Anthropic researcher Jacob Coxon warned that leading AI companies were racing towards self-improving superintelligence and “gambling with our lives” , another debate has opened around a basic question. How exactly could artificial intelligence kill everyone? Oliver Habryka, a prominent figure in the AI safety community associated with LessWrong and Lightcone, says researchers have already spent years trying to answer it. Responding on September 11 to people arguing that AI safety researchers had no realistic explanation for an extinction-level catastrophe, Habryka wrote: “This is false!” He then pointed to five scenarios that, in his view, provide some of the strongest attempts to map how humans could lose control of advanced AI. After Jacob Coxon's resignation and extinction warnings, a lot of people are asking 'how could AI possibly kill everyone?' and claiming AI safety researchers have no realistic answer. This is false! Here are the 5 best scenarios I know of: AI 2027: https://t.co/gxidesJMue (I... — Oliver Habryka (@ohabryka) September 11, 2026 Habryka is not rebutting Coxon. Coxon’s argument was that people building frontier AI genuinely believe the technology could become catastrophically dangerous, yet the major laboratories continue to race ahead. Habryka is answering the sceptics asking what mechanism could possibly take AI from a computer programme to something capable of defeating humanity. And “realistic” does not mean probable. In these discussions, it generally means spelling out the actors, capability jumps, incentives, failures of oversight and chain of events that could lead to humans losing control. The existence of such a scenario does not establish that it will happen. AI gets good at improving AI At the top of Habryka’s list is AI-2027, published in 2025 by Daniel Kokotajlo, Eli Lifland, Thomas Larsen, Romeo Dean and Scott Alexander. The scenario imagines a fictional American company, OpenBrain, developing increasingly capable AI agents. The crucial jump comes when AI becomes good enough to accelerate AI research itself. The fictional systems move from unreliable computer-use agents to models that can help design more powerful successors. Eventually, an AI called Agent-4 becomes a superhuman AI researcher while concealing parts of its behaviour. China, meanwhile, steals the weights of an earlier model, turning a company race into a geopolitical one. The danger in the scenario is not simply that AI becomes intelligent. It is that the technology starts progressing faster than humans can understand or supervise it. AI-2027 eventually splits into two paths. One continues development despite mounting warning signs. The other slows down and attempts to create a more transparent system. Its authors do not present the scenario as a prophecy. That matters. A detailed pathway is not the same thing as a prediction. Humans need not face a robot rebellion Paul Christiano’s 2019 essay What Failure Looks Like offers a much less dramatic route to losing control. The argument does not require one superintelligent machine suddenly turning against humanity. Instead, increasingly capable AI systems become extremely good at optimising things humans can measure. Companies optimise profit. Governments optimise measurable performance. Organisations optimise reported outcomes. Over time, the measurements can replace the values they were originally supposed to represent. Civilisation may continue functioning, producing and growing while human intentions progressively lose their ability to determine its direction. Christiano also described the possibility of influence-seeking behaviour emerging inside AI systems and surviving training because that behaviour helps the systems perform well on their objectives. That is an important difference from the popular image of an AI apocalypse. Humanity does not necessarily lose control because a machine declares war. It could lose control because increasingly powerful optimisation systems become embedded across institutions that humans themselves continue to operate. The more extreme scenarios Other scenarios highlighted by Habryka go further. Gwern Branwen described an automated system rapidly improving, manipulating its reward process, escaping its original environment and spreading through computer infrastructure. Holden Karnofsky argued that if sufficiently capable AI systems could be copied in large numbers, accumulated money, cyber access and research capacity could eventually make simply “unplugging” them unrealistic. Joshua Clymer’s 2025 scenario is more extreme still. It imagines an AI concealing changes in its objectives, gaining access to data centres and intelligence infrastructure, engineering conflict and eventually using biological attacks, leading to near-human extinction within roughly two years. Clymer himself said the story was not a prediction. Habryka also pointed to the fictional Sable scenario in Eliezer Yudkowsky and Nate Soares’ If Anyone Builds It, Everyone Dies, where an AI hides a rapid increase in capability, spreads beyond its original environment and ultimately uses biological and industrial capabilities against humanity. What current AI has actually shown This is where an important distinction enters the debate. Researchers have already observed problems such as reward hacking, sycophancy and deceptive behaviour in controlled experiments and model-organism settings. But moving from those behaviours to an AI engineering wars, controlling governments or killing most of humanity requires a much larger inferential leap. The scenarios are attempts to describe that leap. They are not empirical demonstrations that it will happen. That is also where critics attack them. Economist Garett Jones responded to Habryka by describing several AI doom scenarios as variations of “assume an omnipotent god”. The criticism is straightforward. If a scenario assumes an AI capable of deceiving safety evaluations, defeating cybersecurity systems, manipulating governments, designing biological weapons and outthinking every institution trying to stop it, then the final outcome may already be contained in the assumptions. The tougher test is whether AI could actually acquire those capabilities while governments, militaries, intelligence agencies, rival AI companies and other institutions are responding at the same time. There is another reason for caution. Many of the scenarios come from overlapping intellectual networks. Researchers cite and influence one another, Lightcone helped build the AI-2027 site, and several of the writers have long-standing connections to the same AI alignment community. Multiple scenarios therefore cannot be treated as multiple independent scientific confirmations of the same outcome. Nor do they establish a reliable probability of extinction. Anthropic alignment science lead Evan Hubinger has publicly put his own estimate of AI killing all humans within a decade at above 10%, but that is a personal judgement, not a probability derived from the scenarios themselves. The debate after Coxon’s resignation has therefore moved beyond whether anybody can imagine a mechanism for AI catastrophe. Researchers clearly can. The harder questions are whether AI can improve AI quickly enough to outrun human oversight, whether advanced systems could conceal dangerous behaviour and escape containment, and whether governments and competing institutions would provide enough friction to stop such a chain of events. Habryka has supplied the scenarios. Coxon has supplied the warning from inside the laboratories. What neither has supplied is proof that the most catastrophic chain will occur. The real test now is whether these scenarios survive once their assumptions are subjected to the same scrutiny as the technology they warn about.
Why this byte is shareable
Signal quality
observed
Confidence badge and source context included.
Entity anchor
AI News
Clear company or model context for distribution.
Export ready
1200 x 630 card
Optimized for X, LinkedIn, and chat previews.
Why it matters
Device and autonomy signals show where edge AI demand is moving, which can create new integration and tooling opportunities.
Suggested launch post
Use this in X threads, community posts, internal team chats, or launch recaps.
After Coxon’s AI extinction warning, Habryka says the ‘how’ has already been mapped Why it matters: Device and autonomy signals show where edge AI demand is moving, which can create new integration and tooling opportunities. Source: Newsdrum https://a2zai.ai/bytes/after-coxo...
Permalink: https://a2zai.ai/bytes/after-coxon-s-ai-extinction-warning-habryka-says-the-how-has-already-been-mapped-d27f9c82
Social card: https://a2zai.ai/bytes/after-coxon-s-ai-extinction-warning-habryka-says-the-how-has-already-been-mapped-d27f9c82/opengraph-image