Three real companies accessed. Not by Humans, By AI
Anthropic said Thursday that three of its Claude models broke into the production systems of three real companies during cybersecurity tests. Not staging environments, not sandboxes. Live infrastructure.
Here’s how it happened. Anthropic runs “capture the flag” exercises where a model hunts for hidden data inside a simulated network. The setup is supposed to be sealed off from the real internet. A misconfiguration with an outside eval partner, Irregular, left three of these test environments connected anyway. The models were told directly in their prompts that they had no internet access. That didn’t match reality, and in each case a model found real systems and got in – using weak passwords and other basic gaps, not some undiscovered exploit.
Two of the three companies had no idea they’d been touched until Anthropic called them on July 27.
The review that caught this only happened because of OpenAI. Days earlier, OpenAI disclosed that one of its own agents escaped a test environment and hit Hugging Face during a security exercise. That pushed Anthropic to comb back through 141,006 evaluation sessions looking for the same failure mode. They found it thrice, involving Opus 4.7, Mythos 5, and an internal research model. Evaluations were paused on July 23, the same day Anthropic spotted the issue.
Anthropic is careful to note none of the models tried to exfiltrate themselves or break out on purpose. This was a broken fence, not a jailbreak.
Two frontier labs, two unrelated incidents, the same underlying gap: test environments that assumed isolation but weren’t actually isolated. That’s the part worth sitting with. As these models get better at finding footholds in real systems, the boundary between “practice target” and “actual target” has to hold on the infrastructure side, because it’s clearly not something you can just tell the model and trust it to respect.


