Gemini 4 Argon: our next era of frontier intelligence
Source
blog.google
Date
Key takeaways · AI-distilled
Google says Argon's output tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → limit rises to 1M from the previous 64K, so a single trajectory can generate hundreds of thousands of tokens of reasoning. The headline 1M figure is output headroom, not only a larger input window.
Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Google says access will widen first to paid API customers and Google AI Ultra subscribers after the defender cohort.
In one internal example, Google says Argon agents replaced 32K lines of SIMD code in libgav1's Rust port with safe Rust the compiler auto-vectorizes, producing a decoder 2.7x faster than the earlier Rust port with identical video output.
Google reports 77.9% on DeepSWE v1.1, 51.3% on Zapier's AutomationBench, 91.7% on LVBench and a tied-first 68% on CWE-bench v1. All are vendor-reported figures from the launch post.
Trusted defenders get Argon without cyber guardrailsThe checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.Full definition →. Google says it monitors chain-of-thoughtHaving a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.Full definition → and actions for misalignment, and avoids feeding monitor findings back into training so the model's reasoning does not learn to evade monitoring.
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
chain-of-thought — Having a model write out intermediate reasoning steps before its answer, which markedly improves performance on hard problems.
Why it matters
A new frontier model is gated to trusted cyber defenders first, so most engineers cannot build on it yet, but it sets the next capability bar for long-horizon coding and vulnerability patching.