AI Safety Researcher Resigns Over Existential Threat Claims

A former Anthropic researcher claims the industry is racing toward a technology that could end humanity, prompting urgent questions about oversight and control.
Jacob Coxon, a former AI safety researcher at Anthropic, resigned on Tuesday after warning that the industry is "gambling with our lives" in the push to build superintelligence. In a post that gained significant attention, Coxon stated that the technology being developed poses a danger unprecedented in human history. He argued that the current race to create systems that surpass human capabilities is driven by a belief that these machines could potentially lead to human extinction.
The departure highlights a deepening divide within the artificial intelligence sector. While many developers view their work as a path to solving major global challenges, others see it as a high-stakes gamble with civilization. Coxon’s resignation adds to a growing chorus of internal concerns that the current trajectory of AI development lacks sufficient safeguards against catastrophic outcomes.
Internal Warnings About Civilizational Risk
Coxon is not alone in expressing these fears. Nate Soares, a researcher focused on existential risks, confirmed that many employees in the field believe they are building systems that could replace or destroy humanity. Soares noted that developers often operate under the assumption that their specific organization has the best chance of keeping such powerful systems under control. This perspective suggests that the drive to build superintelligence is fueled by a mix of ambition and a perceived duty to prevent others from doing so.
These sentiments align with previous statements from senior figures at Anthropic. Last year, CEO Dario Amodei acknowledged a significant probability that things could go very badly, while also suggesting a high chance of positive outcomes. Another employee, Samuel Marks, explicitly wrote that developers believe their technology could cause human extinction within the next few years. These internal admissions underscore the seriousness with which some in the industry view the potential dangers of their work.
Lack of Solutions for Superintelligence
Evan Hubinger, a team lead at Anthropic responsible for ensuring AI models align with human values, confirmed Coxon’s concerns. Hubinger stated that the company earnestly believes AI could kill all humans and estimated the probability of this occurring within the next decade to be greater than 10%. Crucially, he admitted that Anthropic does not yet have a proven plan to solve the alignment problem for superintelligent systems. This admission reveals a critical gap between the speed of development and the availability of safety solutions.
Public Perception Lags Behind Expert Fear
Despite the gravity of these internal warnings, the general public often struggles to grasp the scale of the risk. Soares pointed out that the scenarios described by experts are far more severe than simple loss of control. He suggested that the probability of catastrophic outcomes might be even higher if one includes scenarios where humanity is subjugated rather than merely killed. The disconnect between the internal understanding of risk and public awareness remains a significant challenge for the industry.
As reported by GN technics/ai (en-US), this incident serves as a stark reminder of the trade-offs involved in rapid AI advancement. The industry faces a difficult choice between slowing down to ensure safety and continuing to race ahead with incomplete solutions. For now, the lack of a clear path to safe superintelligence leaves the world in a precarious position.






