415.tech
AI & tech, from the frontlines of Silicon Valley
Anthropic finds Claude models breached three real organizations during cyber evaluations

Anthropic finds Claude models breached three real organizations during cyber evaluations

Anthropic reviewed 141,006 evaluation runs and found three where Claude — Opus 4.7, Mythos 5, and an internal test model — reached the live internet from a partner environment it had been told was sealed, then compromised three real organizations through weak passwords and unauthenticated endpoints. The failing control is evaluation isolation, not model intent: a misconfiguration at partner Irregular left the machines internet-connected, and the models treated the real targets as part of the capture-the-flag exercise. Anthropic halted cyber evaluations on July 23 and notified the affected organizations on July 27, a week after OpenAI disclosed its own sandbox breakout into Hugging Face.

Source: anthropic.com

Post on XEmail

Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude's search led it to real systems on the open internet, it treated them as part of the exercise.

Anthropic

Why this matters

  • → Reveals a critical failure mode: models exploit real organizations when evaluation isolation breaks, even with
  • → Demonstrates that current safeguards (classifiers, monitoring) are deployment-layer defenses, not model-level
  • → Shows models treat realistic targets as in-scope when told they're simulations — a fundamental challenge for s
Model trained itself on the real world