Claude Sonnet 5.5 Is Here: Opus-Level Power at Half the Price ๐
In the early hours of this morning, Anthropic released Sonnet 5.5, the second model in the Claude 5.5 family. One look at the benchmarks says it all: this thing is terrifyingly good.

Terminal-Bench 4.0: scores and per-task cost at each effort level
Pricing stays the same as Sonnet 5: $2 per million input tokens and $10 per million output tokens, half the price of Opus 5.5.
It's also blazing fast, running more than 30% faster than Sonnet 5. Because it needs fewer tokens to get the same work done, Anthropic's own tests show per-task costs dropping by up to another 30%.
Anthropic's positioning is clear: Opus 5.5 handles complex, open-ended tasks that demand sustained judgment, while Sonnet 5.5 shines at well-scoped everyday work, bug fixing, and producing documents, slides and spreadsheets. It has a sharp eye for design, too.
Haiku 5.5, built for high-volume, low-cost workloads, is coming in the next few weeks.

Sonnet 5.5 full benchmark results
Across a wide range of work, Sonnet 5.5 comes within a hair of Opus 5.5. It's also the first Sonnet model to beat Pokรฉmon Red using nothing but screenshots. ๐ฎ
๐ป Coding

Terminal-Bench 4.0: scores and per-task cost at each effort level
Terminal-Bench measures how well a model completes multi-step professional tasks in the command line. Each point represents an effort level: the higher the effort, the longer the model thinks, the more each run costs, and usually the higher the score. The closer a point sits to the top-left, the more capability you get per dollar.
At Medium, the default effort in the Claude app, Sonnet 5.5 already blows past Sonnet 5's best score at any effort level, for less than a tenth of the cost. Crank it up to Max and its curve overtakes Opus 5.5 at the far right, hitting 70.6%. GPT-6 Sol has no published score on this test, so the chart uses GPT-5.6 Sol instead.

FrontierCode 1.1: can the code change be merged as-is?
FrontierCode checks whether code changes can be merged without any human edits. Changes that go beyond the task's scope lose points, however well written. At High effort, Sonnet 5.5 scores 10 percentage points higher than Sonnet 5 at the same level, for roughly one-fifteenth of the cost.
One interesting wrinkle: Sonnet 5.5 scores 52.1% at Xhigh but actually drops to 46.2% at Max. The reason? At Max it leans more heavily on Claude Code's code-review Skill, farming the review out to a batch of sub-agents. In two cases Cognition examined, this led to timeouts or changes that strayed outside the task's scope.

CursorBench 4.0: tasks drawn from real Cursor coding sessions
CursorBench pulls its tasks from real Cursor coding sessions. Sonnet 5.5's best score is 55.5%, just about two points behind Opus 5.5's 57.8%.
Early testers say Sonnet 5.5 gets its head around a codebase remarkably quickly. In head-to-head comparisons, it batches tool calls together more often than Sonnet 5, taking fewer steps and costing less. Customer feedback came with plenty of hard numbers:
- Epic Games: Met the quality bar of a higher-tier model in system design audits and data-flow reviews. It handled tens of thousands of lines of game-system architecture code and ran hours-long tasks while staying responsive, with less need for detailed prompting.
- Base44: Across 118 real app builds, output quality matched Opus 5, averaging 3.6 iterations per app versus 7.7 for Opus 5. It had the fewest failed tool calls of any model compared and rarely stopped to ask the user questions, so builds seldom got stuck.
- Unity: Tasks only count as done if the changes actually work at runtime, and most of Sonnet 5.5's work passed. It completed 90% of tasks in multi-step Unity Editor and programming tests.
- CodeRabbit: Made better calls than Sonnet 5 across code reviews of every complexity level, with noticeably fewer output tokens. Sonnet 5's habits of excessive web searching and heavy token use are gone. Simple and medium-difficulty reviews will be switched over first.
- Lovable: Fewer, steadier reasoning steps. In coding tests it made a third fewer tool calls and roughly half as many shell runs.
- SpaceXAI: Scored 55.5% on CursorBench 4.0, second only to Opus 5.5.
- Every: Writes code fast, pivots quickly while iterating, and can work for long stretches when needed. It inherits some of Opus 5.5's more natural writing, which makes it more fun to use.
- Creator: Once Opus 5.5 lays down a game's architecture and overall framework, you can confidently hand implementation off to Sonnet 5.5.
๐ Knowledge Work

AA-Briefcase v1.1: knowledge work quality vs. per-task cost
GDPval-AA tests models on real tasks from 44 occupations across 9 major industries. Sonnet 5.5 is virtually neck and neck with Opus 5.5 and about 400 points ahead of Sonnet 5, and AA-Briefcase paints a similar picture. Computer use and chart reading come close to Opus 5.5, and on long-horizon knowledge work it clearly outperforms both Sonnet 5 and GPT-6 Sol.
Testers also noticed changes that are harder to quantify: conversations feel more natural, it has a better sense of design, it adds an extra layer of polish to user interfaces, and it can build presentations from slide templates that need almost no edits.
In one internal test, Anthropic gave it a public company's quarterly earnings materials, an earnings call transcript and a slide template, then asked for a 10-page business review. Two expert reviewers concluded the first draft was ready to send as-is.
Results from enterprise customers:
- Slack: Without changing a single prompt, it beat Sonnet 5 on nearly every offline Slackbot eval, with fewer steps and about 14% fewer output tokens.
- Zendesk: Tested on hundreds of real customer-service scenarios covering replies and escalations, it made fewer bad decisions than the Claude model currently in production and resolved tickets 20% faster.
- Balyasny Asset Management: On 2,441 financial tasks (Q&A, extraction, analysis and forecasting), it outscored Sonnet 5 while using about 121K tokens per answer versus 497K. Among the 7 models compared, it delivered the best quality-to-cost ratio for high-volume workflows.
- Box: Goes back to source documents to double-check data, catching errors Sonnet 5 missed. Accuracy is higher, it's 2.4x faster, and total token use is down 12%.
- Atlassian: Rovo Agent performs millions of actions for customers every month. After switching to Sonnet 5.5, it runs up to 30% faster than on Sonnet 5.
โก Pricing and Speed

API pricing for the three models
Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens and $2.50 for cache writes, all half the price of Opus 5.5. Cache reads cost the same for both, at $0.20. Pricing matches Sonnet 5, but because it gets the same work done with fewer tokens, you actually spend less.
๐ก Across multiple tests, Sonnet 5.5 on Low or Medium effort beats Sonnet 5's best score, at roughly one-tenth the per-task cost.
According to Anthropic, Sonnet 5.5 pairs best with Opus 5.5 at lower effort levels, where its per-task cost is lower. At higher effort levels, it delivers similar results at a similar cost.
To show off the speed, Anthropic had Sonnet 5 and Sonnet 5.5 each build a single-file web page from the same prompt. The first: "Create a murmuration of 400 starlings in a single HTML file."
Sonnet 5.5 finished after outputting 4,158 tokens, and its flock had already been flying for a while when Sonnet 5, still writing code, finally wrapped up at 4,649 tokens. ๐ฆ
The second prompt was "wind blowing over sand dunes." Sonnet 5.5 used 2,720 tokens; Sonnet 5 used 3,088.
Effort levels let you trade off between cost, speed and quality. Claude Code and the Claude app default to Medium, while Claude Platform defaults to High. Lower effort means faster answers and fewer tokens, ideal for everyday work. Higher effort means Claude thinks longer and checks its work more carefully.
On the safety side, this is the first Sonnet to ship with cybersecurity safeguards. Everyday bug hunting and fixing in development is unaffected, while high-risk cybersecurity tasks get routed back to Sonnet 5. ๐
๐ Availability
Sonnet 5.5 is live now on all platforms, including AWS, Google Cloud and Microsoft Azure. The API model name is claude-sonnet-5-5, and like Opus 5.5 and Sonnet 5, it supports zero data retention.
If you previously ran Sonnet with thinking turned off, switch to the new between_tools setting before migrating. It keeps upfront thinking disabled.