Press Esc to close

Short

Google ships a Gemma 4 update, sort of

Google is rolling out updates to Gemma 4, and the headline claims are solid on paper: uniform Flash Attention 4 support on NVIDIA Hopper GPUs, with prefill throughput up 25–70% and time-to-first-token dropping as much as 31%. The release also patches the chat template and tool-calling issues, and adds a vision token budget option — bump max_soft_tokens to 1120 for sharper OCR and full 2.51MP detail.

But the comments under the announcement on X are where it gets interesting. Several people, including Petri Kuittinen (with screenshots), say only the README and chat_template.jinja files actually changed — the safetensors weights are identical to three months ago. I wonder why Google shipped this as Gemma 4 instead of 4.1, considering the new claims.

Anyway, the updated collection is live on Hugging Face.

Comments