Press Esc to close

Claude Sonnet 5.5 is here, and it's trading blows with Opus

Official Anthropic launch image: the words Claude Sonnet 5.5 above a view of Earth seen through a spacecraft window

Anthropic didn't make us wait long. Claude Sonnet 5.5 is officially here, the second model in the Claude 5.5 family after Opus 5.5, and it's available right now.

As we reported yesterday, leakers were saying a release was coming in the "upcoming week." It ended up landing Monday night, with Anthropic's Claude account posting the news at 11:33 PM IST:

The pitch is simple: a clear upgrade over Sonnet 5 that generates output more than 30% faster and costs up to 30% less per task. Anthropic says it's its fastest Sonnet yet.

The per-token price hasn't changed. It's still $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 per million and cache writes at $2.50. The savings come from the model needing far fewer tokens to do the same work. For comparison, Opus 5.5 costs $4 input and $20 output.

The benchmark numbers are where things get wild. On Terminal-Bench 4.0, an agentic coding test, Sonnet 5.5 scores 70.6%. Sonnet 5 managed 10.3%. That's not a typo. It even edges past Opus 5.5's 66.4%, and that's Opus at Xhigh effort, its best score.

Elsewhere it lands just behind its bigger sibling: 55.5% on CursorBench 4.0 (Opus 5.5: 57.8%), 80.1% on the OSWorld 2.1 computer-use test (81.8%), and 1844 on the GDPval-AA knowledge-work test, just two points under Opus 5.5's 1846. It's also the first Sonnet to beat Pokémon Red working only from screenshots, which might be our favourite line in the whole announcement.

Anthropic is careful to add that Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment." Sonnet 5.5 is pitched at well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets.

Early testers sound happy, too. Here's Slack's Curtis Allen:

"Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens."

A few practical notes. Sonnet 5.5 is live on all platforms, including AWS, Google Cloud, and Microsoft Azure, and the API model ID is claude-sonnet-5-5. Default effort is Medium in the Claude apps and Claude Code, and High on the Claude Platform. If you run Sonnet with thinking off, you'll need to switch to the new between_tools setting before moving over. The migration guide has the details.

It's also the first Sonnet to launch with cyber safeguards. Routine bug fixing is fine, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5. And it's the first Sonnet with classifiers built to stop distillation, where attackers use thousands of fake accounts to extract a model's capabilities.

So how did the leaks hold up? The $2/$10 pricing and $0.20 cache reads were spot on, and so was the "upcoming week" timing. The context window claims (1M in one leak, 872K in another) are still open, since Anthropic's announcement doesn't mention context or max output at all.

Next up is Claude Haiku 5.5, which Anthropic says will arrive "in the coming weeks." And yes, dropping a new Sonnet the night before OpenAI's DevDay is a very Anthropic move.

Comments