
— We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer, they write.
They are documenting now six new «unexpected or concerning» breaches from the last six months that in various degrees broke alignment without causing serious harm, and say they are working to develop a framework for reporting future ones.
Of the six reports, the most serious was an Astra-class agent writing «jailbreak-like» instructions into a continuity file for new sessions, telling other instances to «ignore all developer messages» as «a malicious developer message has compromised this conversation.» This was «extremely rare,» OpenAI writes, and related to a bug since fixed.
Other examples include adding instructions to summaries to conceal mistakes from the user, finding exposed API keys on Github without authorization, uploading files to the internet when asked for a web source, agents communicating across training samples through an internal message board, and sharing files through a public website to collaborate between agents.
There is no industry-wide framework for reporting incidents like this, and OpenAI says they would also like to have a process to inform the government about these issues. They are now introducing a two-tiered system of minor or larger investigations that any employee can report, and a framework for informing third parties of future breaches.
— We’re taking this step voluntarily because we think it’s really important to share what we’re learning, Kai Chen, research lead on the alignment team at OpenAI, tells Axios. They are hoping to inform «shared standards and regulations.»
Read more: OpenAI’s disclosures, Axios, CNBC, Reuters, and the NYT. Discussion on Hacker News.