Press Esc to close

Google says Gemini 4 isn't underperforming — some Googlers reportedly disagree

Google's official Gemini 4 Argon key art: the Gemini spark logo and the words Gemini 4 Argon in white, with a large, softly blurred number 4 over a blue gradient background.

Image: Google

Google has been telling anyone who would listen that Gemini 4 Argon is its most capable model yet. Now the framing has shifted a little. According to a Bloomberg report, employees inside Google have found that the model’s real-world performance doesn’t always match the benchmark scores the company put on stage.

The specific complaint is not that Argon is bad. It’s that it does less well when employees actually put it to work, and that it struggles to handle certain coding tasks. That’s a meaningful distinction, because benchmark leadership is exactly what Google led with. 9to5Google reported that during the unveiling, Google showed Argon outpacing competitors including OpenAI’s GPT-6 Astra across many key areas.

Google pushed back firmly on the characterization. Speaking to Bloomberg, the company said it “would be inaccurate to say that Gemini 4 is underperforming in areas such as coding.” It also pointed back to comments from Koray Kavukcuoglu, who runs Google DeepMind, from late September, when he said “it’s a certainty that we are always gonna be at the frontier.”

What makes the story more interesting than a simple model-is-worse-than-advertised headline is that Google’s own staff don’t agree with each other. The report cites one Googler who describes a “large consensus” internally that Argon sits at the frontier, alongside others who worry the model will still lag behind what Anthropic and OpenAI ship next. Both things can be true at once, and that tension is probably the most honest description of where Argon actually stands.

Google bar chart of DeepSWE v1.1 scores, higher is better: Gemini 4 Argon 77.9%, Claude Opus 5.5 74.2%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%.

Google's published DeepSWE v1.1 results for Gemini 4 Argon. Image: Google

The published numbers are genuinely strong. On DeepSWE v1.1, Argon scores 77.9%, ahead of Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. That’s a real lead, not a rounding error, and it’s the kind of result that justifies the frontier language in the announcement. We covered the launch in full when it happened, including the 1 million token output limit and the pricing structure.

Argon’s coding results aren’t limited to that one benchmark either. On CWE-bench v1, which measures how well a model can remediate security vulnerabilities, Argon ties for first with a top score of 68%. The company also claims leading performance in financial research, legal drafting, end-to-end business automation and long video understanding, where it posts 91.7% on LVBench. Those aren’t the kinds of numbers a model gets by accident.

And there’s a counterpoint worth weighing against the skepticism: Google is already using Argon for work that would be very expensive to get wrong. A team of Argon agents analysed fleet-wide profiling telemetry and identified memory optimisations across Google’s data centres, freeing up more than 300 TiB of memory with an estimated total saving of up to 1 PiB. Separately, Argon agents are migrating C and C++ codebases to Rust, scaling from libraries like re2 up to the 800,000-line Fuchsia OS Zircon kernel. For libgav1, agents replaced 32,000 lines of SIMD code and produced a memory-safe video decoder that runs 2.7 times faster than the existing Rust port with identical output.

If Argon were broadly struggling with coding, that last example would be hard to explain. Which is why the more plausible reading of the Bloomberg piece isn’t that the benchmarks are fake, but that benchmarks and production work measure different things. A model can top an eval and still need more hand-holding than engineers expect when the task is ambiguous, the context is messy, and nobody has written the test yet. That gap between scoring well and being useful is a problem the entire industry has, not just Google.

It’s also a good reminder to verify rather than trust. We wrote about Google’s PageBreak work, where the interesting part wasn’t the 500-odd XSS bugs the system found, it was the decision to deterministically validate every one of the model’s claims instead of taking its word. The same instinct applies here. Whether Argon is “the frontier” or “struggling” depends almost entirely on which tasks you hand it, and the only reliable way to find out is to test it against your own work.

For now, Gemini 4 Argon is only available to trusted testers and cyber defenders through the Fairwind Program, with a wider release starting with Google AI Ultra subscribers and paid API customers. Introductory pricing sits at $2 per million input tokens and $10 per million output tokens, with cached input billed at 95% off, before it settles at the standard $4 and $20. Google says it has improved its ability to monitor the model’s internal activations for misuse, along with what it calls its most prompt-injection-resistant model yet.

For now, the honest position is that both the enthusiasm and the scepticism are coming from people with access to the same model. Google will settle the argument the way it always does, by shipping it to everyone and letting the results speak.

Comments