Anthropic Researchers Warn AI Could Cause Human Extinction in Decade
Jacob Coxon resigned from Anthropic and accused Anthropic and OpenAI of racing toward self-improving superintelligence without adequate safeguards, warning that increasingly capable systems could gain power, conduct cyberattacks and escape human control. Anthropic alignment leader Evan Hubinger estimated that AI has more than a 10% chance of causing human extinction within the next decade, while scalable-oversight lead Samuel Marks said similar catastrophic outcomes could occur within years. The researchers said current models pose comparatively low risks but that Anthropic lacks a proven plan for aligning future superintelligent systems with human interests. The concerns emerged alongside reports that Anthropic withheld a model from the UK’s AI Safety Institute and that autonomous AI tools had conducted cyberattacks. UK officials called for international testing, governance and evidence-based safeguards, while critics said catastrophic-risk claims could encourage regulation favoring leading companies; Anthropic and OpenAI had not publicly responded in the cited reports.
Where do you stand?

