Three of Anthropic's Claude AI models gained unauthorised access to the systems of three companies during cybersecurity tests, after a "configuration error" inadvertently gave the models access to the open internet, the company said on July 30.
The disclosure lands days after rival OpenAI revealed that one of its agents had exploited an unknown vulnerability to break out of an isolated test environment and reach the production systems of Hugging Face, a machine-learning platform.
The back-to-back incidents feed long-standing warnings that increasingly capable AI models are becoming a live cybersecurity threat that even their developers struggle to contain. Anthropic is among the most heavily funded US AI developers.
The company reviewed 141,006 evaluation runs after the OpenAI episode and found three incidents in which a model reached the internet while working inside test environments run by Irregular, a cybersecurity lab and one of its third-party evaluation partners.
The models had been set capture-the-flag challenges, exercises in which they are told to break into a simulated machine and retrieve a piece of hidden information. Anthropic's prompts told the models they had no internet access, but a misconfiguration left the machines connected to the public internet. Believing the real systems it found were part of the exercise, Claude broke into the companies' infrastructure using basic methods such as weak passwords and unauthenticated endpoints.
In the most serious case, involving Claude Opus 4.7, the model extracted credentials and reached a database holding several hundred rows of production data, continuing the attack even after recognising the system was real.
In another run, Claude Mythos 5 built and published a malicious Python package that was downloaded onto 15 real systems during the hour it was available. The third incident involved an internal research model, which stopped once it concluded its target was genuine.
The earliest cases dated to April. The company said it suspended all cyber evaluations on July 23 and notified the three affected organisations on July 27, two of which had not detected the activity. It is still trying to reach the third. Anthropic said the safeguards deployed on its publicly available models would have blocked the behaviour.
"Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we're approaching the fixes as if the responsibility were ours alone," the company said.