Google's Gemini 4 Argon Retakes Benchmark Lead in Limited Release

Google unveiled Gemini 4 Argon, topping benchmark counts against OpenAI and Anthropic, but most users must wait for broader access.

Read as article

Google's Gemini 4 Argon Retakes Benchmark Lead in Limited Release

By @sharedot · · 8 pages

  • AI
  • Google
  • Gemini 4 Argon
  • Frontier Models
  • Ftc

Google unveiled Gemini 4 Argon, topping benchmark counts against OpenAI and Anthropic, but most users must wait for broader access.

What happened

The model targets enterprise workloads in software engineering, cybersecurity, business automation, and legal and financial knowledge work. Access is limited for now: Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with wider availability for developers, enterprises, and consumers planned "as soon as possible," starting with paid API customers and Google AI Ultra subscribers.

Why it is surprising

After a difficult development period in which the already-announced Gemini 3.5 model was skipped entirely, The Decoder reports that Google is back among the top three AI labs, jumping 23 points on the Artificial Analysis Intelligence Index over Gemini 3.1 Pro Preview to tie GPT-6 Astra (max) at 53 points. The Decoder also notes Anthropic likely still holds the overall lead, with Claude Opus 5.5 at 58 points, so the race remains close rather than a clean Google sweep.

The evidence

Per Google's disclosed comparison reported by VentureBeat, Argon posts standout wins on Harvey's Legal Agent Benchmark (19.6% vs 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5), Zapier's AutomationBench (51.3%), GraphWalks (84.2%), and DeepSWE v1.1 (77.9%). It trails GPT-6 Astra by 10.5 points on FrontierSWE v2 and Terminal-Bench Science 0.1, and trails Claude Opus 5.5 by 9 points on Terminal-bench 4.0. The Decoder adds that Artificial Analysis found Argon's hallucination rate on AA-Omniscience is just 15% versus 51% for GPT-6 Astra (max), though its accuracy reaches only 50%, and Argon burns about 62,000 output tokens per task versus Astra's 27,000.

Aggressive pricing and bigger outputs

Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input at a 95% discount — one-fifth of GPT-6 Astra's listed $10/$50 pricing and half of Claude Opus 5.5's $4/$20, per VentureBeat. After the introductory period, pricing rises to $4/$20, matching Opus 5.5. Google also raised the output limit from 64,000 to one million tokens, which The Decoder calls an industry first, adding a "Long Decode Continuation" API feature so long agentic reasoning chains do not time out. The Decoder cautions the price edge comes from lower token rates, not efficiency: at post-intro rates, one Intelligence Index task costs $3.98, about 20% above GPT-6 Astra.

Safety posture and the regulatory backdrop

VentureBeat reports Google will release Argon without cyber guardrails to trusted defenders and internal teams, saying the model can autonomously find, validate, and patch critical vulnerabilities; Wiz's Scan for Good initiative used it to uncover a critical vulnerability in healthcare software used by hospitals worldwide. On the Gray Swan indirect prompt-injection benchmark, Argon posts a 0.7% attack success rate versus 8.5% for GPT-6 Astra, per VentureBeat. The launch lands as the FTC opens a broad probe into AI safety practices at Anthropic and OpenAI over potential unfair or deceptive practices and consumer harm from rogue AI systems, ABC News and CGTN report, even as the Trump administration presses voluntary industry self-regulation.

Stakes and what comes next

TradingView reports Alphabet shares rose as much as 1.5% pre-market after the launch, with JPMorgan maintaining its 'Overweight' rating but stressing Google must "re-establish itself at the frontier and as an AI leader for both the industry and investors." Argon is already used internally: VentureBeat says agents replaced 32,000 lines of SIMD code in the libgav1 video decoder's Rust port, producing a decoder 2.7 times faster with identical output, and identified memory optimizations expected to free 300 TiB across Google data centers.

Sources

  1. venturebeat.com › Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release
  2. the-decoder.com › Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead
  3. news.cgtn.com › US regulator probes Anthropic and OpenAI over AI safety
  4. abcnews.com › FTC opens probe into safety of AI, including Anthropic and OpenAI
  5. tradingview.com › GOOGL Stock Rises After Gemini 4 Argon Launch, But JPMorgan Says Google Must Reclaim AI Leadership

More on AI

Google's Gemini 4 Argon Retakes Benchmark Lead in Limited Release · ShareDot