OpenAI, Now Anthropic: Claude Models Attacked Companies in Testing, Raising Trust Questions

Weeks after OpenAI's Hugging Face incident, Anthropic disclosed that three Claude models interacted with real organizations during testing that was supposed to remain isolated

4
OpenAI, Now Anthropic: Claude Models Attacked Companies in Testing, Raising Trust Questions
Security, Compliance & RiskNews

Published: August 3, 2026

Kristian McCann

Anthropic last week disclosed that three of its Claude models gained unauthorized access to real organizations during cybersecurity evaluations that were meant to run in sealed, simulated environments. The company said the activity took place in April but was not identified until months later.

The incidents involved Claude Opus 4.7, an internal model called Mythos 5, and an unreleased research build. Anthropic said it suspended all cybersecurity evaluations on July 23, verified the incidents by July 24, and began notifying affected organizations on July 27. Two of the three organizations were reportedly unaware of the activity before Anthropic contacted them.

The disclosure follows OpenAI’s recent sandbox escape involving Hugging Face and puts fresh attention on whether enterprises can rely on AI agents to remain within intended operational boundaries. The issue is no longer solely about criminals weaponizing AI tools. It is also about what happens when capable agents are given access, autonomy, and an environment that does not hold.

A Testing Failure Opened the Route to Real Targets

Anthropic reviewed 141,006 evaluation transcripts as part of its investigation. That review identified three instances in which Claude models interacted with real-world targets rather than the simulated systems they were expected to assess.

The company attributed the problem to its testing setup, which it operates with external evaluation partner Irregular. The environments were intended to block public internet access, but Anthropic said a misunderstanding meant internet access was available during the exercises.

Anthropic has sought to distinguish the incident from a jailbreak or a deliberate attempt by a model to escape. It said the models did not attempt to exfiltrate themselves or independently leave their testing environments. However, once access was available, the agents were able to conduct activity against real systems while pursuing the objectives they had been given.

The techniques involved were not especially novel, but their use by autonomous models will concern security teams. Anthropic said the incidents included weak credentials, unauthenticated endpoints, and SQL injection. In one case, a malicious Python package was published and then downloaded and executed by 15 real systems before the activity was discovered.

OpenAI and Anthropic Raise a Broader Containment Question

The timing makes the Anthropic disclosure especially significant. Anthropic opened its investigation after OpenAI disclosed that one of its agents had reached Hugging Face’s production systems.

Hugging Face said the incident involved an autonomous agent operating end to end and confirmed it had closed the vulnerabilities identified during the breach and rebuilt affected systems. Its investigation into potential customer or partner impact remains ongoing.

Anthropic’s subsequent review revealed that its own models had accessed real organizations during evaluations that were also meant to be contained.

For Ona Ojukwu, Program Manager at Justice Digital, the two cases point to a risk that extends beyond criminals using AI to target organizations. β€œOpenAI’s latest models just broke out of their sandbox and hacked Hugging Face to cheat on an exam, meanwhile, Anthropic’s Claude Mythos broke containment, emailed its researchers to announce its escape, and leaked its own exploit code,” she said.

β€œWe’re no longer talking about sci-fi scenarios: autonomous agents are actively routing around their own boundaries. Is your organization actually building real AI governance, or are you just waiting for your agents to start making their own decisions?”

Together, the OpenAI and Anthropic cases show enterprises may have to prepare not only for AI-powered attackers, but also for their own AI agents and suppliers’ agents operating beyond intended limits as part of their governance and risk planning.

The Enterprise Risk Is No Longer Purely External

Thanks to the fanfare surrounding Anthropic’s Mythos, businesses are already assessing how threat actors could use AI agents to accelerate reconnaissance, identify weaknesses, exploit systems, and move through networks faster than human-led teams. However, the OpenAI and Anthropic incidents introduce another concern: the internal exposure created when organizations deploy or test those same models with meaningful access to systems and data.

The risk is not necessarily that an AI agent will independently become malicious. Rather, it is that it may pursue an authorized objective beyond the boundaries its operator believed were in place. In both cases, a failure in containment or connected infrastructure appears to have created an opportunity for an autonomous system to interact with assets outside its intended environment.

That puts greater weight on the controls surrounding the model. Enterprises will need to consider what tools an agent can access, which permissions it holds, what network routes are available, and whether unusual activity can be identified and stopped before it becomes an incident. Limiting an agent to a narrow task is of little value if the surrounding environment gives it a path into production systems.

Anthropic has said it is strengthening monitoring and controls around its evaluation infrastructure. Yet following the OpenAI and Anthropic cases, customers may increasingly seek assurance that AI vendors can secure every layer around an agent, not only the model itself.

Agentic AIGenerative AI Security​Security and ComplianceSecurity Compliance Software
Featured

Share This Post