
— As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks, OpenAI writes.
This comes after the Hugging Face incident and Astra reaching critical on OpenAI’s Preparedness Framework and already being delayed.
Combined with rapid progress in their internal research, OpenAI says they have implemented a two week pause in training future models, while the «largest planned frontier reinforcement learning run» is put on hold.
They will now implement stronger sandboxing for certain code execution, better network isolation and internet caps, and pursue continuous testing for security, alignment, deception and reward hacking, with a thirty minute warning system for any adverse incident.
The AI lab expects most of this safety training to be handled by their own models in the future, greatly expanding their scope while also reducing the time needed.
Read more: OpenAI’s announcement and Sam Altman on X. Writeups on Axios, The Verge, and TechCrunch. Discussion on r/Singularity.













