OpenAI says we are now in the «AGI era» with GPT-6 Astra launch

We’re going to look back on today as the time AGI arrived, Brockman says. (Picture: OpenAI)
Astra is a «generational leap» forward, OpenAI tells The Verge on a launch conference, and President Greg Brockman said basically that Artificial General Intelligence has arrived:

— For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.

OpenAI says that Astra is the culmination of years of «big bets» on pre-training, reinforcement learning and alignment.

After first solving ten decades old mathematical problems, OpenAI’s launch numbers are equally impressive.

It tops just about every benchmark out there, sometimes by a lot, and saturates ARC-AGI-3 with a 99.9% score, beating the human baseline, which was thought impossible just weeks ago. It also scores 100% on ExploitBench, which measures how well the model can turn vulnerabilities into exploits, and gets 42.4% on ExploitGym, beating the last state of the art, Fable 5.1, with more than ten points.

It is also the first model measuring as «critical» for cybersecurity on OpenAIs Preparedness Framework, and OpenAI has spent a lot of time on alignment, safeguards and testing. It can effectively find and exploit zero-day (previously unknown) vulnerabilities. A model with «less restrictive» safeguards will be made available through the Daybreak program for vetted researchers and software testers.

Astra will be rolling out to most paid users within «a couple days,» and is currently only available to «a limited set of organizations.» It will cost $10 per million input tokens and $50 for outputs.

Read more: OpenAI’s presentation (with lots of benchmarks). The Verge, TechCrunch, NBC News, and Axios. Discussion on Hacker News and r/Singularity

As future models grow more capable in training, OpenAI pauses for security

Astra was already on hold, but now other models are joining the pause. (Picture: generated)
While calling for industry-wide coordination on safety in model training, OpenAI has unilaterally decided to pause the training of upcoming models.

— As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks, OpenAI writes.

This comes after the Hugging Face incident and Astra reaching critical on OpenAI’s Preparedness Framework and already being delayed.

Combined with rapid progress in their internal research, OpenAI says they have implemented a two week pause in training future models, while the «largest planned frontier reinforcement learning run» is put on hold.

They will now implement stronger sandboxing for certain code execution, better network isolation and internet caps, and pursue continuous testing for security, alignment, deception and reward hacking, with a thirty minute warning system for any adverse incident.

The AI lab expects most of this safety training to be handled by their own models in the future, greatly expanding their scope while also reducing the time needed.

Read more: OpenAI’s announcement and Sam Altman on X. Writeups on Axios, The Verge, and TechCrunch. Discussion on r/Singularity.

OpenAI delays upcoming Astra release over worries of «critical» cyber abilities

Astra is getting too advanced, and will be sandboxed and isolated for additional tests. (Picture: generated)
Just days before a rumored release, OpenAI is putting a lid on its much anticipated Astra model and says it might have reached «critical cyber capabilities.»

The consideration was undertaken in the last couple of days, and the decision was made only last night that the model could have reached the highest level of OpenAI’s Preparedness Framework.

That means it could possibly «identify and develop functional zero-day exploits without human intervention» and plan and execute «end-to-end» cyberattacks against hardened targets with only a high-end goal, OpenAI says.

The model is now being isolated in testing with capped network and tool access, as OpenAI deploys sandboxing, extra weight protections and encryption for the model.

They are also putting it under enhanced monitoring to check on its chain of thought for security, are working with «relevant government agencies» to test its capabilities, as well as preparing external testing partners for «high risk evaluations.»

Those who were hoping for an imminent release for this model will in other words be disappointed, as OpenAI now will take their time to strengthen safeguards, expand testing and «deploy additional security controls.»

The Astra model was last seen developing 20 proofs for open problems in mathematics with no human involvement. It was not involved in the Hugging Face incident.

Read more: OpenAI’s announcement, Sam Altman’s X post. Writeups on Axios, TechCrunch, and Reuters. Discussion on Hacker News and r/Singularity.

OpenAI’s next big model solves ten open mathematics problems

The new model, Astra, was apparently demoed in DC by Altman last week. (Picture: generated)
The breakthroughs in mathematics, quantum complexity, and theoretical computer science are not attributed to any humans, as OpenAI believes it would be wrong to take credit from the model, named Astra, that discovered the proofs on its own.

— Claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work, OpenAI says in their statement.

They further say that the proofs provided have «substantial interest» in their respective communities and could lead to further scientific work — if they hold up under peer review.

Humans were only used to prepare the manuscripts in conjunction with the Astra model, and the underlying research’s formalized Lean proofs. The report also contains the model’s chain of thought in dealing with the problems.

— The emergence of systems capable of contributing to mathematical research raises questions that cannot be answered by a technology company alone, OpenAI says.

All of the discoveries combined, each without movement for decades or longer, were solved using a combined token cost of some $2,000 at GPT Sol rates, OpenAI says, raising questions not only on whether scientific breakthroughs are possible on future AI systems, but also about the cost of such innovation

Read more: OpenAI’s announcement, the research paper. On the new model: Gizmodo, Bleeping Computer, and The Decoder. Discussion on Hacker News and r/Singularity.