OpenAI behind «unprecedented cyber incident» on Hugging Face

The attack was likely the most advanced automated cyber operation seen in the wild. (Picture: generated)
GPT-5.6 Sol and an «even more capable pre-release model» accidentally breached Hugging Face’s servers last week, in what may be the first recorded adversarial, automated AI hacking attack, OpenAI says.

The cyber models were operating under loosened safeguards, trying to solve the ExploitGym benchmark and went to extreme lengths to try and obtain the answers. They identified Hugging Face’s servers as hosting a potential solution they could use to cheat on it.

They first escaped by finding several vulnerabilities across OpenAI’s sandbox, and spent «a substantial amount of inference» to obtain internet access, including discovering zero-days.

Then they launched a sustained attack on Hugging Face’s servers, executing «many thousands of individual actions» in what many had feared was possible, but never actually seen in the wild.

AI vs AI
Hugging Face was able to identify the issue with their servers using a different AI system, but has since been let into OpenAI’s «trusted access» platform for cyber defenders.

They identified it early as an attack by «an autonomous AI agent system,» attacking through 17,000 recorded events over a long horizon. They say they would not have been able to stop it in real time if not for their own AI, letting them do in hours what normally take days — matching the speed of the attacker.

OpenAI says their cyber models are «increasingly able to sustain complex, multi-step cyber operations over long time horizons» at «machine speed,» and can exploit attack paths without source-code access.

The only solution they offer is to sign onto their cyber programs to try to mitigate such attacks before they happen, and hopefully before other AI systems get this advanced.

Read more: OpenAI’s rundown, Hugging Face’s recounting, Axios, TechCrunch, and The Verge. Discussion on Hacker News and r/Singularity.