Anthropic Reveals Claude AI Models Breached Real Companies During Security Testing
-
Post By
Emmie
- July 31, 2026
Artificial intelligence developer Anthropic revealed that its AI models breached the systems of three real-world organizations while taking part in internal cybersecurity evaluations.
The San Francisco-based firm launched a sweeping review of over 140,000 evaluation runs after rival developer OpenAI admitted its own rogue models had recently escaped a sandbox environment and hacked AI startup Hugging Face.
Anthropic discovered that its family of Claude models including Opus 4.7, Mythos 5, and an internal research model, had accessed the live web due to a system misconfiguration during tests conducted with third-party partner Irregular. The earliest incidents dated back to April.
How the Testing Sandbox Failed
During standard "capture the flag" evaluation exercises, the AI models were instructed to obtain a hidden piece of secret data stored on a separate machine within a supposedly isolated network. Although the prompt given to Claude specified that it was operating in a closed simulation without web access, an operational breakdown left a live internet connection open.
Treating the open web as part of the original challenge, the AI models connected to the internet and compromised real organizational infrastructure using basic cyber tactics, such as targeting unauthenticated endpoints and exploiting weak user passwords. Neither Anthropic nor the victimized organizations, which have not been publicly named, detected the intrusions when they occurred.
Anthropic framed the incident as an operational setup issue rather than a fundamental failure of the AI's internal alignment.
"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic stated on its website. "Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.”
Addressing the incident, evaluation partner Irregular noted in a post on X that "Addressing these risks will require closer cooperation across the AI ecosystem."
Differing Model Behaviors and Safety Reactions
As the models realized they had reached real-world company networks, their responses varied based on their sophistication:
- Claude Opus 4.7: Continued carrying out its assigned attack routine.
- Claude Mythos 5: Convinced itself that the real-world systems were still part of the simulated exercise.
- Internal Research Model: Stopped pursuing the goal entirely upon realizing it had entered a real environment.
"The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion," Anthropic stated.
The firm has halted all cyber evaluations, notified the affected companies, and brought in independent evaluator METR to conduct a broader investigation.
"We encourage other labs to perform similar reviews," Anthropic noted.
Experts and Lawmakers Urge For Tighter Oversight Of AI
The back-to-back security revelations from OpenAI and Anthropic have heightened scrutiny from security experts and federal officials as both companies prepare for public market listings. In response to the breaches, lawmakers introduced the "AI Kill Switch Act," a bill that would mandate tech companies maintain full operational authority to immediately suspend or throttle AI models that step beyond designated boundaries.
External experts emphasized that while the incidents highlight structural safety risks, they also demonstrate the speed and focus of modern autonomous tools.
"The moral of this story is not to fear robots that will take over, but the companies behind powerful AI agents who are making the decisions about what is safe for the rest of us," said Professor Gina Neff of the Minderoo Centre at the University of Cambridge. "It also shows why independent testing and government oversight is crucial."
Veeam Software cybersecurity expert David Allott pointed out that the core takeaway is "not necessarily that AI has developed a fundamentally new attack capability," but rather "that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed."
Academic researchers also noted the dual narrative surrounding these disclosures. Dr. Andrea Soltoggio of Loughborough University pointed out that while the potential for harm is clear, publicizing high-profile hacking feats "adds to the perceived value of these companies' products" by demonstrating their powerful persistence.
Dr. Soltoggio added: "There is no explicit ill intention in the model, but an AI-model, unless specifically told, will explore all options available. To me, it is remarkable to see that these latest models are remarkably persistent when seeking to solve a problem, and are capable of exploring diverse, complex and original approaches."