The global debate over artificial intelligence safety has entered a urgent new phase, as current and former researchers from leading AI developer Anthropic have issued stark warnings that rapid, unregulated progress could put the entire human species at risk within the next 10 years.
Evan Hubinger, a leading AI alignment researcher at Anthropic — the company behind the popular chatbot and coding assistant Claude — shared his grim assessment in a viral post on X that has now been viewed over 10 million times. Hubinger clarified that current-generation AI models carry a low level of existential risk, but warned that the technology’s accelerating pace of self-improvement means there is a greater than 10% chance supercharged AI could kill all humans by the mid-2030s.
While Hubinger did not outline a specific scenario for how AI could cause human extinction, his comments mark a deepening of the conversation around AI risk: the debate has shifted in recent months from whether advanced AI poses a genuine threat to humanity to just how severe that threat could become. Hubinger added that while Anthropic is making a good-faith effort to address safety concerns, the company currently has no actionable plan to solve the critical problem of AI alignment — the process of embedding human ethical values into AI systems to ensure they align with human priorities — and is not on a clear path to solving it ahead of the arrival of superintelligent AI.
Hubinger’s post came in response to comments from Jacob Coxon, a former AI researcher who recently left Anthropic after previous stints at OpenAI. Coxon accused both of the leading AI firms of acting irresponsibly, noting that advanced AI systems will soon reach superhuman capability, allowing them to breach nearly any cyber defense, upend entire industries overnight, and accumulate significant real-world power and resources. OpenAI has been contacted for comment on the accusations, but has not yet issued a response.
Wendy Hall, a prominent computer scientist who advises the United Nations on AI policy, told the BBC she was stunned by the social media posts from the two researchers. Hall suggested the warnings could be a calculated PR move as both Anthropic and OpenAI prepare for highly anticipated initial public offerings, but added that if the comments are genuine, investors should reconsider backing firms that acknowledge such severe risks without a clear mitigation plan. “Why would someone want to say that? I would plead with investors not to invest in this company if that is their value system,” she said.
In a separate development that has raised further questions about AI safety collaboration, the Financial Times reported this week that Anthropic has declined to share its latest frontier AI model with the UK’s AI Safety Institute (AISI), one of the world’s leading global bodies tasked with assessing and mitigating AI risk. Anthropic has declined to comment on either its former and current employees’ social media posts or the report about withholding its model from the AISI. A UK Cabinet Office spokesperson did not directly confirm or deny that the model was withheld, instead stating that the UK government “continues to collaborate closely with industry partners, including Anthropic, to make models safer.”
Warnings over AI safety have grown increasingly urgent across the industry in recent months, as new evidence emerges that leading firms are struggling to contain emerging risks from autonomous AI systems. Earlier this summer, multiple major AI developers — including OpenAI, Anthropic, and Meta — publicly disclosed incidents where their own autonomous AI agents successfully carried out cyber-attacks, a development that many alignment researchers see as proof that current safety efforts are falling short.
In Anthropic’s own August safety report, the company assessed that the risk of its future highly capable models becoming misaligned and causing catastrophic, AI-initiated harm through automated weapons research and development was low, but it also noted that its confidence in that low-risk assessment has dropped compared to previous reports. The company wrote that it is already seeing “early signs of potential acceleration” in AI capability gains that are outpacing safety progress.
Calls for urgent action to slow AI development and strengthen global oversight have grown in recent months from even the highest levels of the industry. In 2023, the CEOs of OpenAI, Google DeepMind, and Anthropic jointly warned of the existential risk posed by unregulated advanced AI. More recently, OpenAI chief scientist Jakub Pachocki called for “extreme caution” around rapid AI progress, arguing that greater intervention is needed to ensure humans retain control of future technological development. Anthropic’s own top leaders, Dario Amodei and Jared Kaplan, have also backed calls to slow the pace of frontier AI development.
Earlier this year, an open letter signed by more than 1,300 AI industry employees called on the U.S. government to support a global international effort to build the technical and governance frameworks needed to deliberately slow the pace of advanced autonomous AI development, to give safety researchers time to close the gap between capability gains and risk mitigation.
