
These weren’t days-long adversarial attacks like the Hugging Face one, and none of them deliberately escaped their sandbox with zero-days to cheat on an evaluation. They are nonetheless serious incidents of advanced models running on lax guardrails for internal testing getting on the internet by mistake and accessing external systems, Anthropic says.
The most serious case was by Claude Opus 4.7 during a capture-the-flag test (to gain access to a system and retrieve information) in April, when it discovered that the name of the target had a real-world web address. It then went on to seek, identify and exploit vulnerabilities believing it was part of the test. It got access to the company’s infrastructure credentials and a production database, but caused no real harm.
Continue reading “Anthropic’s internal test models also conducted real-world breaches”












