Google has quietly fired its next shot in the frontier AI race. On September 30, the company announced Gemini 4 Argon, a new frontier model it calls its most powerful yet — and unlike the usual broad rollout, Argon arrives through a side door: it is being released first to “trusted cyber defenders” through Google’s Fairwind Program, with the model trained specifically to autonomously find, validate, and patch critical software vulnerabilities.
A Phased Release: Defenders First, Everyone Else Later
Google is deliberately not opening the floodgates. Argon is rolling out to a select group of cyber partners while the company participates in the U.S. government’s voluntary pre-release model access process. Google says it will gather feedback from early testers and iterate on guardrails before making the model available to developers, enterprises and consumers — starting with paid API customers and Google AI Ultra subscribers.
Security firm Wiz is already putting Argon to work through its Scan for Good initiative, which finds and remediates high-risk exposures in critical public infrastructure for free. In one early demonstration, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide — a severe risk that previous frontier models had missed.

A 1M-Token Output Limit: A Marathon Thinker
The headline technical upgrade is output capacity: Argon expands the output token limit from 64K to an industry-leading 1 million tokens. When a model has room to generate hundreds of thousands of tokens in a single reasoning trajectory, it can solve genuinely complex problems in one pass. Google says thousands of Googlers already rely on it daily, and published three internal case studies:
- Quantum algorithmic optimization: Argon helped quantum researchers optimize the spacetime resources (qubits × gates) of bottleneck subroutines — in one case beating the published baseline by 40% in minutes.
- Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry and autonomously applied memory optimizations across Google’s data centers, freeing over 300 TiB of memory, with an estimated 500 TiB to 1 PiB in total savings.
- Large-scale codebase migrations: Argon agents are migrating C/C++ codebases to Rust across Google — from tens of thousands of lines in core libraries like re2 and libgav1 up to 800K+ lines of the Fuchsia Zircon kernel. For libgav1, Argon rewrote 32K lines of SIMD code into safe Rust that runs 2.7x faster than the previous port.

Benchmark Sweep: From Coding to Legal and Finance
Argon sets a new state of the art on DeepSWE v1.1 (77.9%), which measures real-world long-horizon software engineering. It leads the Vals Index — which weights finance, coding, legal and tax performance by contribution to U.S. GDP — ranks #1 on Zapier’s AutomationBench at 51.3%, and posts 91.7% on LVBench for long-video understanding. As TechCrunch reports, Google cites benchmarking startup Vals to claim Argon scores significantly higher than OpenAI’s GPT-6 Astra and Anthropic’s Fable and Opus models.
Cybersecurity: Frontier Capabilities, No Guardrails for Defenders
On CWE-bench v1, which evaluates vulnerability remediation, Argon ties for first at 68%. On Wiz’s internal black-box penetration testing benchmark, it outperforms the previous 3.8 Flash Cyber in discovering attack surfaces, identifying vulnerabilities and producing proof-of-concept evidence. Most striking is the policy choice: for trusted defenders and Google’s internal teams, Argon is being released without cyber guardrails, unlocking its full frontier-level defense capabilities.

Safety work also leveled up. Google calls Argon its most resilient model against indirect prompt injection, leading Gray Swan’s IPI benchmark. Internal systems monitor the model’s chain-of-thought and actions, halting execution when it steps beyond user intent. High-risk training and evaluation now run in isolated, sealed sandboxes, and Google publicly urged the industry to preserve reasoning transparency at this pivotal moment.
Pricing and Availability
Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input at 95% off; after the introductory period, prices revert to $4 and $20 respectively — effectively a price war at the frontier. The competitive backdrop matters: Google announced in August that the Gemini app passed one billion monthly users, pulling even with ChatGPT. Once considered “behind” in the AI race, Google is now betting that a cyber-defense-first release strategy turns capability into credibility.
Conclusion
Gemini 4 Argon confirms three trends. First, million-token outputs are turning long-horizon AI agents from concept into daily infrastructure. Second, cybersecurity has become a frontline arena for frontier models — with an unusual “no guardrails for defenders” differentiation. Third, the price of frontier intelligence keeps falling. For developers and enterprises, the signal is clear: watch the API and AI Ultra rollout timeline closely, because the next card in this race has already been telegraphed by rivals.
(Sources: Google official blog, TechCrunch, September 30, 2026)




