September 29, 2026
Claude Sonnet 5.5: Benchmarks, Price and When to Use Opus
Claude Sonnet 5.5 keeps the $2/$10 price and nearly matches Opus 5.5. Independent tests show where it saves money, where Opus is cheaper, and what to test.
Article focus
Anthropic's Claude Sonnet 5.5 keeps Sonnet 5's price, runs faster and lands 2 points behind Opus 5.5 on an independent index. But at its top setting it costs more per task than Opus 5.5. Here's what the numbers show, when each model is the better deal, and how to test before you switch.
Section guide
Claude Sonnet 5.5 is Anthropic's new mid-tier model, released on September 28, 2026. It keeps Sonnet 5's price of $2 and $10 per million tokens, and independent tests put it just 2 points behind Opus 5.5. It's a clear upgrade over Sonnet 5. But it isn't always the cheaper choice: at high settings, Opus 5.5 can deliver more for less.
Key Takeaways
- Claude Sonnet 5.5 launched on September 28, 2026, at the same $2 input and $10 output per million tokens as Sonnet 5.
- Artificial Analysis scores it 56 at max effort, up 18 points from Sonnet 5 and 2 behind Opus 5.5.
- At low, medium and high effort, it beats Sonnet 5 on both score and cost per task, but at max effort it costs more per task than Opus 5.5.
- Opus 5.5 at high effort outscored Sonnet 5.5 at xhigh for less money. Sonnet 5.5 wins on speed and on the cheapest tasks.
- Higher-risk cyber tasks visibly fall back to Sonnet 5. Test first.
What Is Claude Sonnet 5.5?
It's the second model in Anthropic's Claude 5.5 family, after Opus 5.5 a week earlier.
Vendor claim: Anthropic's launch page says Sonnet 5.5 runs more than 30% faster than Sonnet 5. It also says the model "typically needs far fewer tokens to do the same work," and costs up to 30% less per task in Anthropic's testing.
Here are the basics:
| Detail | Claude Sonnet 5.5 |
|---|---|
| Release date | September 28, 2026 |
| Model ID | claude-sonnet-5-5 |
| Price per million tokens | $2 input, $10 output |
| Prompt caching | $0.20 cache reads, $2.50 cache writes |
| Effort settings | Low, medium, high, xhigh, max |
| Where to get it | Claude Platform, AWS, Google Cloud, Microsoft Azure |
The price is the same as Sonnet 5's. Sonnet 5 launched at an introductory $2 and $10, and Anthropic kept that price in August instead of moving to the planned $3 and $15, VentureBeat reported. Zero data retention is available.
Anthropic pitches Sonnet 5.5 as a faster, cheaper partner to Opus 5.5. It lists these strengths:
- Agentic coding and fixing bugs
- Knowledge work, such as polished documents, slides and spreadsheets
- Long-horizon tasks that run over many steps
- Image understanding and design judgment
A smaller Haiku 5.5 is due in the coming weeks, VentureBeat said.
How Does Claude Sonnet 5.5 Score on Benchmarks?
Close to Opus 5.5 on most of Anthropic's chart, and ahead on one test.
Vendor claim: Here are Anthropic's own figures. Settings vary by test, so treat small gaps with care.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% |
| FrontierCode 1.1 | 46.2% | 42.4% | 54.4% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1 (knowledge work) | 1844 | 1449 | 1846 |
| OSWorld 2.1 (computer use) | 80.1% | 57.0% | 81.8% |
| Humanity's Last Exam | 64.5% | 54.9% | 67.7% |
Two footnotes matter. Anthropic says the Terminal-Bench 4.0 score for Opus 5.5 uses xhigh effort. It also says a pre-release bug affected structured outputs on GDPval-AA, though it expects little impact.
Independent data: On the Artificial Analysis Intelligence Index, Sonnet 5.5 scores 56 at max effort. That ranks third of the 216 models it tracks.
Opus 5.5 scores 58, Fable 5.1 scores 53 and Sonnet 5 scores 38. The jump is real. That's 18 points over Sonnet 5, at the same price per token.
Is Claude Sonnet 5.5 Cheaper per Task?
Cheaper than Sonnet 5 at most settings, but not always cheaper than Opus 5.5. Price per token is only half the story. The other half is how many tokens a model uses to finish a job.
Independent data: Here's what Artificial Analysis measured across effort settings. Sonnet 5 and Opus 5.5 figures come from its comparison page.
| Effort | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Max | 56 for $7.60 | 38 for $5.09 | 58 for $5.98 |
| Xhigh | 52 for $2.74 | 34 for $2.87 | 56 for $3.46 |
| High | 47 for $1.08 | 32 for $1.79 | 54 for $1.82 |
| Medium | 41 for $0.59 | 28 for $1.00 | 51 for $1.34 |
| Low | 36 for $0.41 | 24 for $0.51 | 42 for $0.55 |
Each cell shows the index score and the cost per task on the index.
Against Sonnet 5, the story matches Anthropic's claim. At low, medium and high effort, Sonnet 5.5 scores far higher and costs less per task. At high effort, it scores 15 points more for about 40% less.
Max effort is the exception. There, Sonnet 5.5 costs $7.60 per task, more than Sonnet 5's $5.09 and more than Opus 5.5's $5.98. Artificial Analysis said Sonnet 5.5 at max used the most output tokens per task it had seen. The cheaper tokens can't make up for that many of them.
Max effort is also slow to answer. Artificial Analysis measured about 328 seconds before the first answer token at max, because the model thinks first.
When Is Opus 5.5 the Better Deal?
More often than the price tag suggests. Opus 5.5 costs twice as much per token, but it uses fewer tokens at its higher settings.
Three comparisons from the table stand out:
- Opus 5.5 at high beats Sonnet 5.5 at xhigh. It scores 54 against 52 and costs $1.82 per task against $2.74, so it's both better and about a third cheaper.
- Opus 5.5 at max beats Sonnet 5.5 at max. It scores 58 against 56, and costs $5.98 against $7.60.
- Opus 5.5 at low roughly matches Sonnet 5.5 at medium. It scores 42 against 41, for $0.55 against $0.59.
Our view: if your work needs a score in the 50s on this index, Opus 5.5 at medium or high looks like the better buy. Anthropic agrees Opus 5.5 is stronger for hard, open-ended work. Our Claude Opus 5.5 guide covers its costs by setting in detail.
One caveat: this is one index of tests, and your prompts, tools and output lengths will change the math, which is why you should run your own numbers.
How Does It Compare With GPT-6 Sol and Astra?
Well, and at the same list price as OpenAI's closest rival. GPT-6 Sol also costs $2 input and $10 output per million tokens, as we covered in our GPT-6 Sol and Luna guide.
Independent data: On the same Artificial Analysis index, here's how the models line up:
| Model and setting | Index score | Cost per task |
|---|---|---|
| Claude Sonnet 5.5, max | 56 | $7.60 |
| GPT-6 Astra, max | 53 | $3.26 |
| Claude Sonnet 5.5, xhigh | 52 | $2.74 |
| GPT-6 Sol, max | 48 | $1.06 |
| Claude Sonnet 5.5, high | 47 | $1.08 |
| GPT-6 Sol, medium | 40 | $0.25 |
Sonnet 5.5 at high effort and GPT-6 Sol at its top setting are almost a tie, on both score and cost. At the low end, GPT-6 Sol is much cheaper per task. At the top, Sonnet 5.5 at max outscores every OpenAI model on this index, but it pays for that with many more tokens.
On Anthropic's own chart, GPT-6 Sol leads Sonnet 5.5 on FrontierCode, 49.3% to 46.2%. Sonnet 5.5 leads on knowledge work, with 1844 on GDPval-AA against 1487. Our GPT-6 Sol vs Opus 5.5 comparison covers the wider matchup.
Where Does Claude Sonnet 5.5 Win?
On speed, on the cheapest tasks and on some agentic coding work.
- Speed. Artificial Analysis measured Sonnet 5.5 at 85 to 139 output tokens per second across settings. Opus 5.5 ran at 74 to 94. Sonnet 5 ran at 56 to 79.
- The lowest cost per task. At low effort, Sonnet 5.5 scored 36 for $0.41 per task, the cheapest result among these three models.
- Short jobs. When a task needs little thinking, such as classification, extraction, short replies or simple tool calls, a lower price per token wins.
- Terminal work. On Anthropic's chart, Sonnet 5.5 beats Opus 5.5 on Terminal-Bench 4.0, 70.6% to 66.4%.
Vendor claim: Anthropic's customers report gains too. Box said Sonnet 5.5 was 2.4 times faster and used 12% fewer tokens, Zendesk said tickets were processed 20% faster, and Lovable reported about a third fewer tool calls.
These are Anthropic's chosen examples, not independent tests.
There's also a fun one. Anthropic says Sonnet 5.5 is the first Sonnet model to beat Pokémon Red from screenshots alone.
What Changes for Security Work?
Some cyber tasks won't run on it, and its reasoning is harder to copy.
Vendor claim: Anthropic says Sonnet 5.5 ships with cyber safeguards similar to those on Opus 5.5. In its words, users "can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5."
Security teams will soon be able to apply to an expanded Cyber Verification Program for access to more advanced abilities. Biology safeguards are the same as Sonnet 5's.
Sonnet 5.5 is also the first Sonnet model with classifiers that block attempts to extract its reasoning. Anthropic says it "expands preserved thinking, so Claude's thinking cannot be decoupled from the account that created it." It adds that most developers won't notice a change.
What this means for teams: most won't notice. If you build security tools, three steps help:
- Log the serving model. A fallback to Sonnet 5 changes quality and cost mid-task, so record which model answered each request.
- Test your security prompts. Run your scanning, triage and code-review prompts on Sonnet 5.5 before you switch, and note which ones fall back.
- Plan for access. If your work needs the restricted abilities, watch for the Cyber Verification Program to open and apply early.
How Should Teams Test Claude Sonnet 5.5?
Treat it as a new model, even though the price is the same. A short plan:
- Run your own eval set. Use real tasks from your product, with known good answers, and score every model on the same set. Our guide to AI agent evaluation shows how to build one.
- Test several effort levels. Compare Sonnet 5.5 at low, medium and high with Opus 5.5 at low and medium. The winner may not be the model you expect.
- Track tokens, not just price. Log input tokens, output tokens and tool calls for every task, because cost per task is the number that matters.
- Route by task. Quick, well-scoped jobs go to Sonnet 5.5. Hard, open-ended work goes to Opus 5.5.
- Watch for fallbacks. Check the model field in each response.
- Roll out in stages. Start with a slice of traffic, then compare quality, latency and cost per task against your current model for at least a week before moving everything.
This takes a few days, not weeks. It's also the only way to know whether the vendor's "up to 30% less" holds for your work.
Should You Switch From Claude Sonnet 5?
For most teams, yes, once your tests pass. The price per token is the same, and the gains on independent tests are large at every setting.
- Switch first: coding help, support replies, document work and agent steps that run at low to high effort.
- Check before switching: any workload you run at max effort, where cost per task went up on Artificial Analysis's tests.
- Check twice: security tools.
If you already moved hard work to Opus 5.5, keep it there. Compare our Opus 5.5 benchmark breakdown with the figures above before you move anything back. Our earlier Claude Sonnet 5 guide covers the production setup that still applies.
How Van Data Team Helps Teams Choose Between Claude Models
We help teams pick and route models based on their own data, not launch charts. That means eval sets from real tasks, cost-per-task tracking, routing rules by task type and staged rollouts with clear quality gates.
Claude Sonnet 5.5 is a strong, fast model at a good price. Whether it's the right one for each part of your product is a question your own numbers should answer. If you want help running that test, our AI governance guide is a good place to start.
Article FAQ
Questions readers usually ask next.
These short answers clarify the practical follow-up questions that often come after the main article.
Need a similar system?
If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.
Book your free workflow review here.
