Read as article
Inside Google, Doubts That Gemini 4 Argon Can Really Code
By @sharedot · · 8 pages
- AI
- Gemini 4
- Model Releases
Bloomberg reports Google employees privately say Gemini 4 Argon's strong benchmark scores do not hold up on real-world coding tasks.
What happened after the Argon unveiling
Google announced Gemini 4 Argon, its new top-tier model, to a small group of trusted cybersecurity partners and said access would expand after further testing, starting with paid subscribers. Bloomberg (via The Star) reports that while Google touted leading benchmark scores, some insiders with direct access to the effort say those metrics "don't tell the whole story," and that the model struggles with certain coding tasks. Alphabet shares initially rose about 1.7% in extended trading but pared gains after Bloomberg's report on the internal scepticism, per The Star.
Why the doubt is surprising
The scepticism cuts against Google's official line. The company told Bloomberg it would be inaccurate to say Gemini 4 underperforms in coding, pointing instead to comments from Koray Kavukcuoglu, the head of Google DeepMind, who said he was "encouraged by the model's performance" and that "it's a certainty that we are always gonna be at the frontier." An employee familiar with development told Bloomberg there is "large consensus" internally that Gemini 4 is at the frontier, yet other employees reportedly believe Anthropic's Fable and OpenAI's Astra are improving faster and will still lead in some areas.
The 'benchmaxxing' problem
Two people familiar with the model told Bloomberg that Gemini 4 appears affected by "benchmaxxing" — an industry tendency to optimize for benchmark scores rather than real products. Edwin Chen, founder of Surge AI, told Bloomberg that benchmark focus can push labs to build models that write code in a particular language instead of well-designed apps: "An analogy would be, 'Oh yeah, my kid got a really good score on the SAT' — but the SAT doesn't translate into real-world performance. It's an incredibly pernicious problem." Bloomberg's sources also flagged weak front-end design skills and a very large, expensive-to-run model.
The evidence on both sides
Google's press materials argue strongly for the model's capabilities. The Cryptonomist reports Argon set a DeepSWE v1.1 software engineering record of 77.9%, topped Zapier's AutomationBench at 51.3%, and jumped to an industry-leading 1 million-token output limit from 64,000 previously. CNBC's Artificial Analysis Intelligence Index places Gemini 4 behind only Claude Opus 5.5 and Claude Sonnet 5.5. Meanwhile, Bloomberg's anonymous sources say internal evaluations show uneven coding ability, and The Star notes Google insisted rigorous internal tests show no struggle with messy real-world coding tasks.
The stakes for Google
Google badly needs Gemini 4 to succeed: versions underpin nearly every product it sells, from AI answers atop Search to Maps, Gmail and Chrome, each with more than a billion users, Bloomberg reports. The pressure is amplified by a costly setback — Bloomberg Intelligence analyst Mandeep Singh estimated abandoned training runs like the scrapped Gemini 3.5 Pro can cost as much as $400 million, and The Star notes Google abandoned that June-planned release after pledging it at I/O. The turmoil has also reshaped the lab itself, with Demis Hassabis ceding day-to-day DeepMind operations to Kavukcuoglu in August and star researchers including Jeff Dean and Noam Shazeer departing, Bloomberg and ASHARQ AL-AWSAT report.
What comes next for Argon
Broader access has no fixed date. The Cryptonomist also reports introductory pricing of $2 per million input tokens and $10 per million output tokens, with cached input at a 95% discount. Analysts told CNBC the real test will come when businesses can deploy Argon widely in production, with IDC's Tim Law saying "the final proof will be in enterprise production environments once the model is fully released."
Sources
- thestar.com.my › Google grapples with employee scepticism about new Gemini 4
- english.aawsat.com › Google Announces Gemini 4 Flagship AI Model After Months of Delays
- en.cryptonomist.ch › Google Gemini AI Model Launches Argon with Premier Cybersecurity Focus
- cnbc.com › Can Google's new model really catch up to OpenAI and Anthropic at the frontier?