ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Gemini 4 Argon Arrives: 77.9% on DeepSWE, Google's New Flagship Goes to Security Partners First

Gemini 4 Argon Arrives: 77.9% on DeepSWE, Google's New Flagship Goes to Security Partners First

AI information • Admin • • 7 views

Gemini 4 Argon is Google's new flagship model, announced on September 30, 2026, and the first release in the all-new Gemini 4 family. In its official blog post, DeepMind called it the start of "our next era of frontier intelligence" — built for deep reasoning across complex, long-horizon workflows, spanning real-world software engineering, enterprise knowledge work (legal and finance), and cybersecurity defense.

Who gets it first: cybersecurity partners

Argon is initially rolling out only to trusted cyber defenders through the Fairwind Program, while Google takes part in the U.S. government's voluntary pre-release model access process. Paid API customers and Google AI Ultra subscribers are next; developers, enterprises and everyday consumers will have to wait. A phased "security first, everyone later" launch is unusual for a Google flagship — itself a signal that Google isn't comfortable releasing some of this model's capabilities all at once.

The scorecard: first place in coding and knowledge work

  • DeepSWE v1.1 (real-world, long-horizon software engineering): 77.9%, a new record; Claude Opus 5.5 scored 74.2% and GPT-6 Astra 74.1% in the same period.
  • The output token limit jumps from 64K to a full 1 million, so a single run can produce a complete reasoning trajectory hundreds of thousands of tokens long — "solve the hard problem in one go" is the pitch Google keeps repeating.
  • Leading on the Vals Index (economic impact of finance, coding, legal and tax work weighted by GDP contribution); #1 on AutomationBench (Zapier's end-to-end business execution benchmark) at 51.3%; a record 91.7% on LVBench for long-video understanding.
  • Cybersecurity is Argon's most emphasized badge: 68% on CWE-bench v1 for vulnerability remediation, tying for first; on Gray Swan's indirect prompt-injection benchmark the attack success rate is just 0.7%, far ahead of GPT-6 Astra's 8.5% — the most injection-resistant frontier model to date.

The most controversial part: a "guardrail-free" version for defenders

Google says it will give trusted defenders and its own internal teams Argon "without cyber guardrails" so they can use the model's full defensive capability. Cybersecurity firm Wiz is already using it through its Scan for Good initiative to protect critical public infrastructure, and in one demonstration it uncovered a critical vulnerability in healthcare software used by hospitals worldwide — one that previous frontier models had missed. While opening the tap for defenders, Google is also hardening four lines of defense: against misuse (refusing CBRN and other harmful requests), against prompt injection, monitoring chain-of-thought to prevent out-of-bounds execution, and hardening sandbox test environments.

Pricing: selling at the "floor price" first

Argon's introductory API pricing is $2 per million input tokens and $10 per million output, with cached input discounted 95% — identical to the launch prices of GPT-6.1 Sol and Sonnet 5.5, and one-fifth of GPT-6 Astra's list price ($10/$50). After the introductory period it rises to $4/$20. With three labs now pinning flagship-grade pricing to the same line, "flagship performance at commodity prices" has become the industry consensus; the contest is no longer about unit price but about who delivers the lowest cost per completed task.

Why this launch matters so much for Google

Google's frontier cadence visibly slipped over the past year: Gemini 3.5 Pro was delayed and then cancelled outright, the flagship line sat at November 2025's Gemini 3 while OpenAI and Anthropic shipped GPT-6 and a new Claude generation, and DeepMind went through a leadership shakeup and core talent departures. But Argon's internal record is solid: it helped the quantum team optimize spacetime resources of key subroutines by 40%, freed 300 TiB of memory across data centers, and is migrating the 800K-plus-line Fuchsia Zircon kernel from C/C++ to Rust — with a Rust rewrite of libgav1 running 2.7x faster than the original. For Google, Argon is more than a model; it's proof that "we're still at the table." What remains to be seen is whether the benchmark crown converts into lower task costs on real production traffic.

Recommended Tools

More