
The Safety and Security Committee of the OpenAI Foundation’s board has a new member in Paul Christiano, a long-running researcher on alignment and safety — and he comes with a warning for the world:
— If we build superintelligence without more robust alignment I expect we will permanently lose control of it, he writes on X, and says — If that happens then most people could die.
The warning comes on the heels of rapid growth in recursive self-improvement recently, and the debate over whether computer systems and AI might gain the possibility to build on themselves at some point in the future. OpenAI reckons they’ll have autonmous AI researchers around 2028, but warns it may come sooner.
His appointment and X essay come just a day after Anthropic researcher Jacob Coxon resigned over alignment, saying his company was «gambling with our lives» by running head-first toward self-improving AI, going viral in the process.
In a reply to his X post yesterday, Anthropic’s Alignment Science Lead Evan Hubinger said he personally thinks there’s a greater than 10% risk that AI could kill all humans, contributing to the dayslong online debate.
OpenAI officially acknowledges that AI is heading for a new era, saying that: «We’ve reached a new chapter in AI capabilities, and that demands a new chapter for AI policy.» They also add that «this makes it increasingly important that these systems are guided by solid safety, security, alignment, and governance.»
Paul Christiano remains hopeful, ending his X essay by saying that «frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination.»
Read more: Christiano’s X post, OpenAI on him joining, OpenAI calls for urgent action, and Christiano’s Wiki page. More on Axios, TechCrunch, and The FT.