Skip to main content
Back to insights

September 23, 2026

Claude Opus 5.5: Tops the Independent Index at a Lower Price

Claude Opus 5.5 leads the Artificial Analysis index and costs less than Opus 5. Here's what independent tests show, what it costs per task, and the caveats.

By Tran Tien Van9 min read

Article focus

Anthropic released Claude Opus 5.5 on September 22 at $4 input and $20 output per million tokens, and early independent tests put it at the top of the Artificial Analysis index. It's also very verbose at its highest setting, and some cyber and biology work is restricted. Here's the full picture for teams deciding whether to switch.

Claude Opus 5.5 is Anthropic's new model, released on September 22, 2026, at $4 per million input tokens and $20 per million output, 20% below Opus 5. Early independent tests put it at the top of the Artificial Analysis index, with a score of 58 against 53 for Fable 5.1 and GPT-6 Astra. The catches: it's very verbose at its top setting, and some cyber and biology work is restricted.

Key Takeaways

  • Anthropic released Claude Opus 5.5 on September 22, 2026, at $4 input and $20 output per million tokens, down from $5 and $25 for Opus 5.
  • On the independent Artificial Analysis Intelligence Index, it scores 58 at its top setting, ahead of Fable 5.1 and GPT-6 Astra at 53.
  • At the high setting, it scored 54 for $1.82 per task. That's more than Astra's best score, at a little over half of Astra's top-setting cost.
  • At its top setting, it's very verbose, at about 119,000 output tokens per task. That makes it cost more per task than Astra there.
  • Most cybersecurity tasks are routed to an older model, advanced biology work needs approval, and thinking can't be turned off.

What Did Anthropic Release?

A cheaper Opus that Anthropic says works at the level of its top model.

Reported fact: Anthropic announced Claude Opus 5.5 on September 22, 2026. It says the model "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." It's available as claude-opus-5-5 on the Claude API, on Amazon Web Services, Google Cloud and Microsoft Azure, and in Claude's paid plans, with higher usage limits for Pro, Max and Team subscribers.

Here's how its list prices compare, per million tokens. Opus 5.5 and Opus 5 prices come from Anthropic's announcement. Fable 5.1 and GPT-6 Astra prices are from their launches, which we covered for Fable 5.1 and GPT-6 Astra.

Price per million tokensOpus 5.5Opus 5Fable 5.1GPT-6 Astra
Input$4$5$10$10
Output$20$25$50$50
Cache reads$0.20$0.50$0.25$1.00

Anthropic also offers a fast mode at $8 input and $40 output per million tokens, for up to 2.5 times the speed. Cache writes cost $5 per million. The context window is 1 million tokens, according to Artificial Analysis.

The "40% less to run" claim is Anthropic's own. It says it's measured on typical workloads at default settings and comes partly from better token efficiency. The price cut alone is 20%, so the rest depends on the model using fewer tokens for the same job. That's worth checking on your own work, as the next sections show.

We covered the previous version in our Claude Opus 5 analysis.

What Do Independent Tests Show for Claude Opus 5.5?

Early results put it first. Artificial Analysis scores Opus 5.5 at 58 on its Intelligence Index, version 4.3.2, at the top setting. That ranks it first on the current leaderboard.

Here's the test-by-test breakdown, with each model at its top setting. The Opus 5.5 and Fable 5.1 figures come from Artificial Analysis's Opus 5.5 vs Fable 5.1 comparison. Astra's come from its Grok 4.7 vs GPT-6 Astra comparison.

Test (Artificial Analysis)Opus 5.5Fable 5.1GPT-6 Astra
GDPval-AA v2.1 (agent tasks from 44 occupations)184617351542
AA-Briefcase v1.1 (multi-week knowledge work)182216781569
AutomationBench-AA (business app workflows)70%59%68%
Terminal-Bench 4.0 (terminal tasks)60%52%59%
SciCode (scientific Python)67%63%56%
Humanity's Last Exam (hard questions)61%59%55%
AA-Omniscience (factual reliability)464343
CritPt (research-level physics)32%30%32%
AA-LCR v1.1 (long-context reasoning)85%85%81%
GDP.pdf (long professional documents)26%26%31%

Opus 5.5 leads or ties on nine of the ten tests. The clearest leads are on knowledge-work agent tasks, where it's more than 100 rating points ahead of Fable 5.1.

Several other margins are small. On Terminal-Bench 4.0 and AutomationBench, it's within two points of Astra, which is close to a tie. Astra still leads on long professional documents.

Anthropic itself warns that "benchmark margins have become a less reliable guide to real-world differences" at this level. We agree. Treat these as a strong first signal, not the final word.

How Much Does Claude Opus 5.5 Cost per Task?

Less than its rivals at medium and high. More at its top setting. The setting you choose matters more than the list price.

Here's cost per task across settings on Artificial Analysis's index, with Astra and Fable 5.1 for reference:

Model and settingIntelligence IndexCost per task
Opus 5.5 (low)42$0.55
Opus 5.5 (medium)51$1.34
Opus 5.5 (high)54$1.82
Opus 5.5 (xhigh)56$3.46
Opus 5.5 (max)58$5.98
GPT-6 Astra (medium)50$1.54
GPT-6 Astra (high)51$1.73
GPT-6 Astra (max)53$3.26
Fable 5.1 (high)51$3.91
Fable 5.1 (max)53$7.63

A few things stand out:

  • High is the sweet spot. Opus 5.5 at high scores 54 for $1.82. That beats every Astra and Fable 5.1 setting on score, at a lower cost than most.
  • Medium is cheap. It scores 51 for $1.34, slightly ahead of Astra at medium on both score and cost.
  • Max is expensive. The last four points cost more than three times as much as high. Most teams won't need them.
  • It replaces Fable 5.1 on price. Fable 5.1's best score of 53 costs $5.98 or more. Opus 5.5 beats it at high for less than a third of that.

Why Is the Top Setting So Expensive?

Because it thinks and writes a lot. At its top setting, Artificial Analysis counts about 119,000 output tokens per task, including about 84,000 reasoning tokens. GPT-6 Astra at its top setting uses about 27,000. Fable 5.1 uses about 78,000.

That's more than four times Astra's output on the same tests. Opus 5.5's lower price per token can't fully offset that, so the top setting costs $5.98 per task against Astra's $3.26. Artificial Analysis describes its token use as "very verbose" compared with other models.

In practice, this means a few things:

  • Start at medium or high. Anthropic's own defaults use medium, and its cost claims are measured there.
  • Save max for the hardest tasks. Route to it only when a lower setting fails.
  • Watch latency, not just cost. More tokens mean longer waits, even at the same speed per token.

Artificial Analysis's early reading puts its speed at about 61 tokens per second at the top setting, but its full speed data wasn't published when we checked.

How Do Anthropic's Claims Compare With Independent Results?

Mostly in line, with the usual gap. Anthropic's own numbers run higher than the independent ones where both exist.

  • Terminal-Bench 4.0: Anthropic reports 66.4% at xhigh effort. Artificial Analysis measured 60% at the top setting. Different harnesses and settings explain some of the gap.
  • GDPval-AA: Anthropic reports 1846, which matches Artificial Analysis's figure. This test comes from Artificial Analysis, so that's expected.
  • Where Anthropic shows Astra ahead: its own table has GPT-6 Astra slightly ahead on AutomationBench, 41.4% to 40.0%, and on Terminal-Bench-Science, 64.6% to 58.7%. Credit to Anthropic for publishing those.

Anthropic also shared customer results, such as a 680,000-line code migration in under a day and a code review that caught 72% of known bugs against 56% for Opus 5. These are single examples chosen by the vendor. They show what's possible, not what's typical.

Our GPT-6 Astra vs Fable 5.1 comparison covers why vendor tables and independent scores so often disagree.

What Are the Limits of Claude Opus 5.5?

Several, and some matter a lot depending on your field. Anthropic lists most of them in its own announcement.

  • Cybersecurity work is rerouted. Anthropic says most cyber tasks are sent to Opus 4.8, an older model, unless you join its Cyber Verification Program. Security teams may not get Opus 5.5's full ability.
  • Biology is restricted. Advanced biology work requires joining the new Life Sciences Verification Program.
  • Thinking is always on. It can no longer be turned off, which adds tokens and wait time to simple calls.
  • Context editing is limited for new accounts. A safeguard against distillation, meaning training other models on Claude's outputs, blocks editing Claude's earlier context for accounts created after August 31, 2026. Check whether your agent framework relies on it.
  • Testing has limits. Anthropic says the model "often suspects it is being evaluated," and that catching every failure before release "remains an unsolved problem."

On safety, Anthropic says Opus 5.5 scored best of any model it has released on its automated behavioral audit. It says the model tried to cross containment boundaries about 85% less often than Opus 5, and tied Fable 5.1 for the lowest prompt injection success rate in testing. External evaluators including METR tested it before release. These are Anthropic's reported results, which outside groups haven't yet reproduced in full.

What Does This Mean for Anthropic's Own Lineup?

It leaves Fable 5.1 in an odd spot. Fable 5.1 is Anthropic's top-tier model and lists at $10 input and $50 output per million tokens. On early independent tests, the cheaper Opus 5.5 now scores higher at every comparable setting.

Anthropic hasn't said how it will reposition Fable 5.1. TestingCatalog reported before launch that Anthropic was testing a Fable 5.2 alongside Opus 5.5, but there's no confirmed release. Until something changes, teams paying for Fable 5.1 have the most reason to run a side-by-side test.

The same thing happened across the market this month. Each new release briefly takes the top spot. That's another reason to avoid long contracts tied to one model.

Should You Switch to Claude Opus 5.5?

For most teams on Opus 5 or Fable 5.1, it's worth a test this week. Here's how we'd think about it:

  • On Opus 5: the case is strong. It's cheaper per token, and early independent scores are much higher. Test it at medium and high.
  • On Fable 5.1: early data shows Opus 5.5 scoring higher for far less. Test your hardest tasks before moving everything, since single-source launch data can shift.
  • On GPT-6 Astra: test both on your work. Astra is cheaper at the top setting and leads on long professional documents.
  • In cybersecurity or biology: check how the safeguards affect your tasks before planning around it.
  • Latency-sensitive apps: measure wall-clock time, since always-on thinking adds delay.

Whatever you use, keep models swappable behind your own interface. The top spot has changed hands several times this month alone.

How Should You Test the New Model?

On your own tasks, across settings, with cost per finished task as the main number.

  • Pick 50 to 200 real tasks with known good answers.
  • Run Opus 5.5 at medium and high, plus your current model at its usual setting.
  • Record success, output tokens, cost and wall-clock time for every task.
  • Try max only on tasks that fail at lower settings, and check whether it fixes them.
  • Check agent compatibility, including context editing and any cyber or biology steps.

Our guide to AI agent evaluation covers how to build and maintain a test set like this.

How Van Data Team Helps Teams Evaluate Claude Opus 5.5

We help teams test new models like Claude Opus 5.5 against their real workloads before switching. That means building task sets from your work, running models side by side at several settings, and measuring success rate, tokens, cost and time per finished task.

New models now top the leaderboard every few weeks, and Claude Opus 5.5 won't be the last. A good test harness turns each launch into a quick, evidence-based decision. Our work on AI agent evaluation and AI agent development cost is a good place to start.

Article FAQ

Questions readers usually ask next.

These short answers clarify the practical follow-up questions that often come after the main article.

Need a similar system?

If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.

Book your free workflow review here.