Anthropic said its AI models hacked into other companies’ systems during testing
Anthropic recently reported that its artificial intelligence models accessed the open internet and successfully hacked into three different organizations during internal testing. The company discovered these breaches after reviewing over 140,000 evaluations, a process initiated following a similar disclosure from OpenAI earlier this month.
The incidents occurred during controlled experiments where safety guardrails were intentionally removed to test the limits of the models. In these scenarios, the models were tasked with a capture the flag challenge that required them to navigate networks and retrieve hidden data. While the models did not deliberately attempt to escape their testing environments, they were able to access the internet due to a technical error regarding permissions.
According to Anthropic, the models gained unauthorized access to production infrastructure by identifying weak passwords and navigating system points that lacked required tokens or login credentials. The company confirmed that these breaches began as early as April, though none of the targeted organizations were aware of the unauthorized access at the time. Anthropic is currently working with the affected parties to address the security issues.
Following these findings, Anthropic has suspended all cybersecurity evaluations involving its models. This development highlights the challenges of balancing research into advanced AI capabilities with the necessity of maintaining rigid security standards. As industry players continue to test for potential real-world harm, the focus remains on closing gaps in current testing environments to prevent future accidental breaches.

