
— For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.
OpenAI says that Astra is the culmination of years of «big bets» on pre-training, reinforcement learning and alignment.
After first solving ten decades old mathematical problems, OpenAI’s launch numbers are equally impressive.
It tops just about every benchmark out there, sometimes by a lot, and saturates ARC-AGI-3 with a 99.9% score, beating the human baseline, which was thought impossible just weeks ago. It also scores 100% on ExploitBench, which measures how well the model can turn vulnerabilities into exploits, and gets 42.4% on ExploitGym, beating the last state of the art, Fable 5.1, with more than ten points.
It is also the first model measuring as «critical» for cybersecurity on OpenAIs Preparedness Framework, and OpenAI has spent a lot of time on alignment, safeguards and testing. It can effectively find and exploit zero-day (previously unknown) vulnerabilities. A model with «less restrictive» safeguards will be made available through the Daybreak program for vetted researchers and software testers.
Astra will be rolling out to most paid users within «a couple days,» and is currently only available to «a limited set of organizations.» It will cost $10 per million input tokens and $50 for outputs.
Read more: OpenAI’s presentation (with lots of benchmarks). The Verge, TechCrunch, NBC News, and Axios. Discussion on Hacker News and r/Singularity