Google announces Gemini 4 Argon: 1M output tokens, 77.9% on DeepSWE v1.1, and a claimed 2.7x faster Rust decoder for libgav1
Synopsis
Google announced Gemini 4 Argon, a frontier model first rolling out to trusted cyber defenders through its Fairwind Program, built for long-horizon deep reasoning, raising the output token limit from 64K to 1M and reporting 77.9% on DeepSWE v1.1, 91.7% on LVBench, 68% on CWE-bench v1, and 51.3% on AutomationBench, alongside internal cases including a 40% improvement over a published baseline in quantum subroutine optimization, over 300 TiB of memory freed in data centers, and a Rust port of libgav1 that replaced 32K lines of SIMD code and runs 2.7x faster than the prior Rust port with identical video output.
Interpretation
Gemini 4 Argon is positioned as a frontier model for long-horizon, multi-step workflows, with its output token limit expanded from 64K to 1M so it can generate hundreds of thousands of tokens in a single trajectory for deeper reasoning. Compared with the previous 64K limit, this expansion targets tasks that must be completed in one long chain rather than short single-turn exchanges. The claim comes from Google's official announcement and is a vendor-stated capability positioning; the text provides no independent third-party evaluation tied to output length.
On public and industry evaluations, Argon reports leading scores: 77.9% on DeepSWE v1.1, 91.7% on LVBench, 68% on CWE-bench v1, and 51.3% on AutomationBench, and is described as leading on the Vals Index, Vals Finance Agent v2, and Harvey's Legal Agent Benchmark. These numbers span real-world long-horizon software engineering, long video understanding, vulnerability remediation, and end-to-end business execution, extending the capability claim beyond coding alone to enterprise knowledge work and multimodality. All are benchmark scores listed in the announcement and are vendor-reported; the text does not provide sample sizes, evaluation protocols, or confidence intervals.
In internal use, Argon helped quantum computing researchers optimize subroutine spacetime resources, beating a published baseline by 40% in one example within minutes; a team of Argon agents analyzed fleet-wide profiling telemetry and autonomously applied memory optimizations, freeing over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings. These are quantifiable outcomes produced inside real production and research workflows, rather than benchmark scores alone. The figures come from Google's internal case descriptions, without stated baseline provenance, measurement methodology, or statistical uncertainty.
For code migration, Argon agents are migrating C/C++ codebases to Rust across Google, scaling from tens of thousands of lines in core libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel; in the libgav1 case, agents replaced 32K lines of SIMD code, producing safe Rust that the compiler vectorizes automatically, yielding a decoder 2.7x faster than the prior Rust port with identical video output. The case grounds agent capability in a specific open-source project and verifiable performance metric, while noting that large-scale rewrites still undergo automated and manual auditing, emulation testing, and review before production. The performance multiple and line counts are provided by the publisher, with no independent reproduction or benchmark details.
Perspective
The release targets trusted cyber defenders, Google's internal teams, and later paid API customers and Google AI Ultra subscribers, making it a phased rollout rather than general availability. The described capabilities apply to long-horizon software engineering, enterprise knowledge work such as legal and finance, multimodal long-video understanding, and defensive cybersecurity; for readers assessing how frontier models perform on real engineering and security tasks, this material offers a set of comparable metrics and case examples.
All benchmark scores and internal cases in the text are vendor-reported, without sample sizes, evaluation protocols, baseline details, or independent reproduction, so these numbers should be treated as company claims rather than independently verified conclusions. The specific measurement conditions behind the 40% quantum optimization improvement, the 300 TiB to 1 PiB memory savings, and the 2.7x libgav1 speedup are not elaborated. In addition, the model is not yet broadly available, so real-world availability, long-term costs beyond the listed price, and how the safety guardrails perform at wider deployment all remain to be observed.
