Image: Google
Google has a new top model. Gemini 4 Argon, announced on Wednesday US time (early Thursday in India), is what Google DeepMind calls "our new frontier model," built for "complex workflows across coding, enterprise knowledge work, and cybersecurity defense."
Introducing Gemini 4 Argon – our new frontier model.
— Google DeepMind (@GoogleDeepMind) September 30, 2026
It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program. pic.twitter.com/X8acOWJOSF
The catch? You probably can't use it yet. For now, Argon is only rolling out to a small group of trusted testers and cyber defenders through Google DeepMind's Fairwind Program.
What's new
The big change is how much Argon can write at once. Google says its output limit jumps to 1 million tokens, up from 64,000 before. That gives the model room to work through long, multi-step problems "in one go," as Google DeepMind puts it.
Google says thousands of its own employees already use Argon. In one
Image: Google
For office-style work, Google says Argon leads the Vals Index (68.9%), which weights finance, coding, legal and tax tasks by their share of US GDP. It also ranks #1 on Zapier's AutomationBench at 51.3%.
Image: Google
It isn't a clean sweep, though. Google's own table shows Argon trailing GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and Claude Opus 5.5 on Terminal-bench 4.0 (57.4% vs 66.4%). Google ran several Argon tests itself, including DeepSWE, per its methodology notes, so independent results will matter.
Not everyone at Google is convinced
Bloomberg reported that some Google employees are skeptical about how well Gemini 4 performs in key areas such as coding. According to a summary of the paywalled story by TheFly, sources said the model does well on industry benchmarks but less well when employees actually put it to work.
LuminaBench, an AI news and benchmarks account on X, framed it as pushback over "benchmaxxing," or tuning a model to ace tests. Its post says Google disputes this, citing a "large consensus" internally that Gemini 4 is at the frontier. LuminaBench added that, in its own testing, the model is frontier. It didn't share any scores.
🚨 Gemini 4 is reportedly getting pushback for benchmaxxing
— Lumina (@LuminaBench) September 30, 2026
Bloomberg says some employees with direct access think it looks much better on benchmarks than it does in actual use, especially coding 💀
Google dispute this and say there’s “large consensus” internally that Gemini 4… pic.twitter.com/jLOMQxz49S
Built for cyber defenders
Google says Argon can find, check and fix software security flaws on its own. On CWE-bench v1, which tests fixing vulnerabilities, Google says Argon ties for first place at 68%. Its chart shows Grok 4.7 and GPT-6 Astra on 68% too.
Image: Google
That's why access is tight. Trusted defenders get Argon "without cyber guardrails," and Fairwind partners may only give access to their security teams, Google DeepMind's program page says. Security firm Wiz used it to spot a critical flaw in healthcare software used by hospitals worldwide, according to Google.
When you'll get it, and the price
There's no date. Google says it will gather feedback and strengthen safeguards first. It's also taking part in the US government's voluntary pre-release testing. When it opens up to developers, enterprises and consumers, it will start "with paid API customers and Google AI Ultra subscribers."
The price is set, though. Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% off the input price. After that, it rises to $4 and $20. Google hasn't said how long the introductory period lasts. That launch price matches GPT-6.1 Sol, which we covered this week.
For most of us, the wait starts now.
example, Argon agents made an existing Rust version of Google's libgav1 video decoder run 2.7 times faster. In another, they freed up over 300 TiB of memory across Google's data centers.The benchmarks, according to Google
On DeepSWE v1.1, a test of long, real-world coding tasks, Google says Argon sets a new state of the art at 77.9%. That's ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%) in Google's chart.
Google says Argon can find, check and fix software security flaws on its own. On CWE-bench v1, which tests fixing vulnerabilities, Google says Argon ties for first place at 68%. Its chart shows Grok 4.7 and GPT-6 Astra on 68% too.
Image: Google
That's why access is tight. Trusted defenders get Argon "without cyber guardrails," and Fairwind partners may only give access to their security teams, Google DeepMind's program page says. Security firm Wiz used it to spot a critical flaw in healthcare software used by hospitals worldwide, according to Google.
When you'll get it, and the price
There's no date. Google says it will gather feedback and strengthen safeguards first. It's also taking part in the US government's voluntary pre-release testing. When it opens up to developers, enterprises and consumers, it will start "with paid API customers and Google AI Ultra subscribers."
The price is set, though. Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% off the input price. After that, it rises to $4 and $20. Google hasn't said how long the introductory period lasts. That launch price matches GPT-6.1 Sol, which we covered this week.
For most of us, the wait starts now.




Comments