Google delivers Gemini 4 Argon in limited release, nudging the frontier ahead

Argon does push the frontier a little, but benchmarks don’t point to a showstopper. (Picture: Google)
Their first fresh frontier model in nearly a year isn’t getting a public release just yet, it is being rolled out to «trusted cyber defenders» and government to test and harden its capabilities first.

The model is built for «complex, long-horizon workflows,» Google says, and is supposedly excels at «real-world software engineering and enterprise knowledge work.»

It is also good at cybersecurity, being able to find, validate and patch vulnerabilities, and the release to testers will be without cyber guardrails so they can test out these defenses.

Argon also slightly pushes the frontier ahead, scoring marginally better than GPT-6 Astra, Opus 5.5 and Fable 5.1 in most benchmarks Google tested with, and is ahead by quite a bit on Arena.ai’s text benchmark, while coming in eighth on coding.

On the Artificial Analysis benchmark it scores a joint second behind Opus 5.5, shared with Fable 5.1 and Astra. It also hallucinates way less, at a rate of just 15% versus Fable’s 69%.

The real-world testing release is in order to protect against misuse, prompt injections, and misalignment, and to harden systems, Google says. They also say they will be monitoring Argon’s chain-of-thought throughout these «pivotal moments of increased capabilities while navigating alignment risks.» Google signed a voluntary safety agreement between industry and the White House just yesterday.

There seems no timeline for general release, but Google says it will roll out to the API and AI Ultra subscribers first, and Argon is available now to members of the Fairwind program. It costs $2 per million input tokens and $10 per million out.

Read more: Google’s presentation. CNBC, Ars Technica, and VentureBeat. Discussion on Hacker News and r/Singularity.