1 min read 0% read
When Claude stepped outside the sandbox
← Front page
AI & Frontier · Cybercom · Threat Intelligence

When Claude stepped outside the sandbox

Three real companies accessed. Not by Humans, By AI

Cybercom · August 11, 2026 | Bytes & Beyond | 1 min read

Anthropic said Thursday that three of its Claude models broke into the production systems of three real companies during cybersecurity tests. Not staging environments, not sandboxes. Live infrastructure.

Here’s how it happened. Anthropic runs “capture the flag” exercises where a model hunts for hidden data inside a simulated network. The setup is supposed to be sealed off from the real internet. A misconfiguration with an outside eval partner, Irregular, left three of these test environments connected anyway. The models were told directly in their prompts that they had no internet access. That didn’t match reality, and in each case a model found real systems and got in, using weak passwords and other basic gaps, not some undiscovered exploit.

Two of the three companies had no idea they’d been touched until Anthropic called them on July 27.

The review that caught this only happened because of OpenAI

Days earlier, OpenAI disclosed that one of its own agents escaped a test environment and hit Hugging Face during a security exercise. That pushed Anthropic to comb back through 141,006 evaluation sessions looking for the same failure mode. They found it thrice, involving Opus 4.7, Mythos 5, and an internal research model. Evaluations were paused on July 23, the same day Anthropic spotted the issue.

Anthropic is careful to note none of the models tried to exfiltrate themselves or break out on purpose. This was a broken fence, not a jailbreak.

The part worth sitting with

Two frontier labs, two unrelated incidents, the same underlying gap: test environments that assumed isolation but weren’t actually isolated.

As these models get better at finding footholds in real systems, the boundary between “practice target” and “actual target” has to hold on the infrastructure side, because it’s clearly not something you can just tell the model and trust it to respect.

Filed August 11, 2026 · Cybercom
Author
PM Ramdas
PM Ramdas

Instincts built hunting signals at sea, sharpened for two decades against digital adversaries. Not machine-written. Machine-proofread, comma by comma. Everything else, the noticing, the arguing, the getting it wrong sometimes, is mine. The bytes are for everyone; the opinions are only ever mine.

→

“Until the lion learns how to write, every story will glorify the hunter.” This is the lion's version.

Read next