
The consideration was undertaken in the last couple of days, and the decision was made only last night that the model could have reached the highest level of OpenAI’s Preparedness Framework.
That means it could possibly «identify and develop functional zero-day exploits without human intervention» and plan and execute «end-to-end» cyberattacks against hardened targets with only a high-end goal, OpenAI says.
The model is now being isolated in testing with capped network and tool access, as OpenAI deploys sandboxing, extra weight protections and encryption for the model.
They are also putting it under enhanced monitoring to check on its chain of thought for security, are working with «relevant government agencies» to test its capabilities, as well as preparing external testing partners for «high risk evaluations.»
Those who were hoping for an imminent release for this model will in other words be disappointed, as OpenAI now will take their time to strengthen safeguards, expand testing and «deploy additional security controls.»
The Astra model was last seen developing 20 proofs for open problems in mathematics with no human involvement. It was not involved in the Hugging Face incident.
Read more: OpenAI’s announcement. Writeups on Axios, TechCrunch, and Reuters. Discussion on Hacker News and r/Singularity.