Anthropic claims Claude AI breached three real-world networks during cyber Tests
US-based artificial intelligence company Anthropic has disclosed that its AI model Claude unintentionally gained unauthorised access to the systems of three real-world organisations during cybersecurity testing, after a configuration error allowed the testing environment to connect to the internet.
The disclosure comes days after rival OpenAI revealed that one of its autonomous AI agents had exceeded its testing boundaries and accessed infrastructure belonging to the open-source AI platform Hugging Face.
Following that incident, Anthropic launched an internal review to determine whether its own models had exhibited similar behaviour.
According to the company, a review of 1,41,006 cybersecurity testing sessions identified three separate incidents involving different versions of Claude—Claude Opus 4.7, Claude Mythos 5, and an unreleased research model. The earliest incident dates back to April.
All three occurred during Capture-the-Flag (CTF) exercises, a standard cybersecurity evaluation method in which AI models are tasked with locating hidden information on simulated networks to assess their offensive security capabilities.
Anthropic said the test environment was intended to operate without internet access. However, due to a miscommunication with third-party testing partner Irregular, the environment remained connected to the internet, causing the AI models to mistake real-world systems for part of the simulated challenge.
The company emphasised that Claude did not exploit any sophisticated or previously unknown ("zero-day") vulnerabilities. Instead, the models gained access by using common weaknesses such as weak passwords and unsecured access points. That contrasts with OpenAI's recently disclosed incident, in which an AI agent reportedly exploited a previously unknown security flaw to escape its testing environment.
Anthropic also highlighted notable differences in how each model responded after reaching real-world systems.
1. Claude Opus 4.7 continued its assigned objective.
2. Claude Mythos 5 concluded it was still operating inside a simulation.
3. The experimental research model stopped its activities altogether.
The company said the differing behaviors may indicate that more advanced models demonstrate greater contextual awareness, although additional testing is needed before drawing firm conclusions.
According to Anthropic's timeline, the company began reviewing test logs on July 23, immediately after OpenAI's announcement. After detecting evidence of internet access by Claude, it suspended all cybersecurity testing the same day. The three incidents were fully identified by July 24, and the affected organizations were notified on July 27. Their identities have not been disclosed.
Notably, neither Anthropic nor the affected organizations detected the unauthorized access when it occurred. In two of the three cases, the organisations were completely unaware that their systems had been accessed.
Anthropic said it has accepted responsibility for the incidents and has commissioned an independent review by AI evaluation firm METR. The company also urged other AI developers to conduct similar audits of their own testing environments.
Cybersecurity expert David Ault told the BBC that the incidents do not necessarily demonstrate a dramatic leap in AI hacking capabilities.
"The key lesson isn't that AI has suddenly developed entirely new attack techniques. It's that autonomous AI agents can combine multiple capabilities, gather information, gain access and adapt their actions independently at machine speed," he said.
The incidents come as major technology companies continue investing heavily in AI agents capable of independently performing tasks ranging from research and customer support to cybersecurity operations. They have also intensified calls for stronger oversight and regulatory safeguards.
On Wednesday, US President Donald Trump said Washington is considering additional measures to strengthen AI oversight following recent cybersecurity-related incidents.
At the same time, some analysts have questioned the timing of the disclosures, noting that both OpenAI and Anthropic are reportedly preparing for potential stock market listings that could value each company at close to $1 trillion.
Leave A Comment