A Resignation Sparks Industry Alarm

A senior safety researcher at Anthropic, Evan Hubinger, stated on Tuesday that artificial intelligence holds a probability greater than 10% of causing human extinction within the next decade. This public concession followed a resignation thread from Jacob Coxon, a former researcher at both Anthropic and OpenAI. Coxon left his post earlier this week after publicly accusing top AI firms of operating without regard for long-term safety. He argued that these labs are currently fixated on a race toward self-improving superintelligence. They are ignoring the massive risks tied to building systems that might eventually act against human interests.

Coxon noted that internal culture at these companies often favors speed over caution. He claimed that researchers remain stuck in a cycle of competition. They fear that if they slow down, another less careful company will reach the finish line first. This environment prevents any single lab from pausing their work. Consequently, the industry continues to push forward with recursive self-improvement research. This work is moving faster than most observers predicted just two years ago.

The Reality of Recursive Self-Improvement

Hubinger, who serves as the alignment science lead at Anthropic, confirmed that he shares many of these concerns. While he emphasized that current models pose a low risk, he highlighted the danger of future, autonomous systems. These future machines could possess the capability to rewrite their own code. Such systems might escape human oversight while gaining access to critical power and physical resources. Hubinger admitted that the company lacks a plan to ensure such systems remain safe.

This lack of a clear safety plan keeps internal researchers awake at night. They are building tools that may one day perform at a superhuman level. These systems will likely possess the speed to revolutionize major global industries in a single day. Yet, the mechanisms to control these systems remain hypothetical. The current pace of advancement leaves little room for rigorous safety testing. Experts fear this oversight could result in a catastrophe once these models begin to act on their own.

Future Implications for Global AI Policy

The warnings from both Coxon and Hubinger reflect a growing divide between technical progress and safety preparedness. Coxon proposed that a global pause on specific model improvements might be the only way to prevent a disaster. He believes that without an international agreement, individual labs will never stop their pursuit of dominance. His suggestion calls for a temporary moratorium on increasing the capabilities of the most advanced models.

Such a move would be unprecedented in the technology sector. Leaders in Washington and the European Union have debated regulation for years, but no consensus exists on how to handle recursive self-improvement. The broader situation remains dire. Researchers like Coxon are now calling for a shift in perspective across the entire field. They argue that developers must prioritize the long-term safety of humanity over the immediate thrill of engineering breakthroughs. If the current trajectory continues, the window to ensure these technologies remain friendly to humans may close sooner than expected.