Google launched Argon, the first model of the Gemini 4 generation, late on September 30. Three things stand out: the single-response output limit jumps from the previous 64K to 1M tokens; it scores 77.9% on DeepSWE v1.1; and most counterintuitively, the first users are not developers or subscribers but trusted cyber defenders in the Fairwind Program.

Strong on long-horizon work, weaknesses in plain sight

Per the official announcement, Argon targets complex, long-horizon workflows: real-world software engineering, legal and financial knowledge work, and cyber defense. The 1M figure is an output ceiling, not a context window — Google publishes no context length for Argon at all. On benchmarks, Google says the 77.9% on DeepSWE v1.1 is a new state of the art; its own comparison chart puts Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. It also leads the Vals Index (68.9%), Zapier's AutomationBench (51.3%, rank #1), and long-video understanding LVBench (91.7%), and ties for first on CWE-bench v1 at 68% alongside Grok 4.7 and GPT-6 Astra.

But a third-party comparison table shows where it loses: 55.0% on FrontierSWE v2 versus Astra's 65.5%, and 57.4% on Terminal-Bench 4.0 versus Opus 5.5's 66.4%. One outlet flatly noted that all four headline numbers come from Google's own benchmark submissions. Winning long-horizon tasks while losing short ones is consistent with the 1M-output positioning.

Internal wins: quantum, memory, C++ to Rust

The official blog showcases three internal cases: helping quantum researchers optimize qubit-times-gates resources of subroutines, beating a published baseline by 40% in minutes; a team of Argon agents analyzing fleet-wide profiling telemetry to autonomously apply memory optimizations, freeing over 300 TiB with an estimated 500 TiB to 1 PiB total savings; and migrating C/C++ to Rust — 32K lines of SIMD code in libgav1 were replaced with auto-vectorizable safe Rust, running 2.7x faster than the prior Rust port, with an 800K+ line Fuchsia Zircon kernel migration underway. Note these are all Google's own accounts with no external replication yet.

The safety narrative is the real story of this launch

Argon was trained to autonomously find, validate, and patch vulnerabilities. Wiz is already using it in its Scan for Good initiative to scan critical public infrastructure for free; in an early demonstration it uncovered a critical vulnerability in healthcare software used by hospitals worldwide that previous frontier models had missed. For trusted defenders and internal teams, Google explicitly releases Argon without cyber guardrails to unlock full defensive capability. On the protection side: internal-activation monitoring, leading prompt-injection robustness on Gray Swan's IPI benchmark, and misalignment monitoring over chain-of-thought and actions. Google is engaged in the U.S. government's voluntary pre-release evaluation process; paid API and AI Ultra subscribers come later. Introductory pricing is $2/$10 per million tokens, reverting to $4/$20 — from half of Opus 5.5's price to parity.

One detail worth watching: OpenAI shipped GPT-6.1 Sol at DevDay, and a day later Google seized the safety narrative with "white hats first, guardrails off." When the output ceiling hits 1M tokens, the real question is no longer whether the model can think, but whether you dare let it think unsupervised in one go.