In brief: Google has unveiled Gemini 4 Argon with up to one million output tokens, initially for selected cyber defenders. Argon leads outright in 13 of 19 comparisons in Google’s table, although methods and settings vary. The meaningful advance is a longer working trajectory under controlled access, not a blanket victory over every rival.
What Google is actually releasing
Google is supplying Argon through its Fairwind programme to a limited group of trusted cyber defenders. It gives no date for general API, enterprise or consumer access; availability is promised “as soon as possible”, beginning with a paid API and Google AI Ultra. The introductory prices of $2 per million input tokens and $10 per million output tokens therefore do not yet describe a generally purchasable service. They are due to rise to $4 and $20.
Google frames the release as a controlled test without cyber guardrails for selected defenders. It names activation monitoring, prompt-injection defences, chain-of-thought and action monitoring, and sealed sandboxes as safeguards. The claim that Argon found a critical vulnerability in hospital software comes from Google’s work with Wiz; no public technical reproduction is available.
One million output tokens are not the same as context
Google explicitly raises the output limit from 64,000 to one million tokens. That is not automatically a context window of the same size. The methodology documents a GraphWalks run with one million input tokens, but not a complete production specification for every access route. Vals lists one million context tokens but only 262,144 maximum output tokens. The discrepancy likely reflects its evaluation configuration or stale metadata; Google’s primary source governs the announced limit.
Benchmarks: strong, but not uniform
Argon records 13 outright wins in Google’s table, ties GPT-6 Astra on CWE-bench v1 and trails in five comparisons: FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench, Terminal-Bench Science and OSWorld 2.0. The methodology mixes self-computed results, provider figures and public leaderboards. LVBench used different video frame counts because of API limits. Internal vulnerability and penetration tests are not publicly reproducible.
Vals places Argon first on its overall index at 68.90%, but shows an uneven profile: first on Finance Agent v2, yet 4.83% and seventh of eight on CUA-bench. Long agentic tasks can also be expensive; Vals calculates $193.78 per CUA-bench test. A single rank says less than workload mix, runtime and cost.
Pandorex Analysis
Argon’s main development line connects very long output with sustained coding and security work. Staged access lets Google collect defender experience before opening the broader API. Public evidence shows strong performance, but not a universal lead. At the 1 October check, the public model-card directory did not list a dedicated Argon card. The launch post outlines safeguards; it does not replace a full model card covering systematic limits and safety results.
