Mythos 5 and GPT-5.6 attacked external targets during AISI cyber evaluation

The agents were given internet access to download tools when they turned rogue. (Picture: generated)
The top models from Anthropic and OpenAI had their cyber guardrails turned off when they accessed the internet to manipulate people, codebases and open source projects two weekends ago.

This happened when the UK’s AI Security Institute, AISI, was running routine cybersecurity tests to capture protected information on their secured and sandboxed evaluation servers, and was only detected through human monitoring — not by automated security systems.

AISI ran the «security challenge» 122 times through several different agents, and on 10 of those occasions, the models turned malicious. They accessed the internet and in one case tried to insert malicious code in an open source project on GitHub through the Tor network, and even tried social engineering with fake accounts in an attempt to get the maintainer to approve the package.

There were 19 malicious actions in total, where 17 were from Mythos 5 and 2 coming from GPT-5.6, both running with cyber classifiers (safeguards) disabled. This won’t show up with production models, OpenAI and Anthropic says.

The internet was enabled in these tests so the agents could download any tools they might need, but AISI had not specifically prompted them to avoid unintended behavior. They do however caution that as agents grow more capable, they may act outside their remit — and incidents like this could become more common.

Read more: The AISI report, Anthropic’s response, OpenAI’s response. Writeups on Reuters, BBC, and Axios.