September 23, 2026
Claude Opus 5.5 Benchmarks: What the Chart Shows and Hides
Claude Opus 5.5 benchmarks vs Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol, row by row, with the footnotes explained and independent results side by side.
Article focus
Anthropic's launch chart shows Claude Opus 5.5 leading seven of nine rows against Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol. The footnotes tell you how much of that to trust. Here's the chart row by row, what each footnote changes, and how independent tests compare.
Section guide
Claude Opus 5.5 benchmarks from Anthropic's launch chart show it leading seven of nine rows against Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol. Astra leads the other two. But the footnotes matter: settings differ, most rival numbers come from other companies' reports, and several leads sit inside the noise. Independent tests confirm Opus 5.5 is ahead overall, by a smaller margin on some rows.
Key Takeaways
- Anthropic's chart shows Claude Opus 5.5 leading seven of nine rows. GPT-6 Astra leads on business workflows and agentic science research.
- In the six rows with an Astra number, Opus 5.5 leads four. Three rows have no Astra number, and two have no OpenAI number at all.
- Several leads are within noise, including FrontierCode by 1.1 points and Chartography by 0.6.
- On Terminal-Bench 4.0, Anthropic's chart shows an 8.5-point lead over Astra. Artificial Analysis's own run shows 1 point.
- Independently, Opus 5.5 scores 58 on the Artificial Analysis index, against 53 for Fable 5.1 and Astra, 51 for Opus 5 and 47 for GPT-5.6 Sol.
What Does Anthropic's Claude Opus 5.5 Benchmarks Chart Show?
A broad lead, with two exceptions. Here's the chart from Anthropic's Opus 5.5 announcement, reproduced row by row. A dash means Anthropic didn't list a number.
| Benchmark (Anthropic's chart) | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (agentic coding) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 (agentic coding) | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam, with tools | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 (science research) | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0, partial (computer use) | 81.8% | 80.7% | 74.0% | — | — |
| Chartography, with tools (chart reading) | 89.0% | 88.4% | 83.4% | — | — |
Opus 5.5 is top on seven rows. GPT-6 Astra is top on two. Credit where due: Anthropic highlighted Astra's two wins on its own chart rather than leaving them out.
Look closer at the dashes, though. Three rows, CursorBench, OSWorld and Chartography, have no GPT-6 Astra number. Two of them have no OpenAI number at all. On those rows, Opus 5.5 is only being compared with Anthropic's own models.
What Do the Footnotes Change?
A lot. The footnotes under the chart explain how the numbers were produced, and they affect how you should read them.
- Opus 5.5 runs at its highest effort. Unless noted, its results use max effort. That's the most expensive setting, and not what most teams use day to day.
- Terminal-Bench compares different settings. Opus 5.5 is shown at xhigh effort and GPT-6 Astra at high effort, "as reported by OpenAI." Anthropic says these are each model's highest scores.
- Rival numbers come from other sources. GPT-6 Astra and GPT-5.6 Sol figures on Terminal-Bench are as reported by OpenAI. AutomationBench results were run and reported by Zapier.
- Safeguards swap in older models. When safeguards stepped in, cybersecurity tasks were finished by Claude Opus 4.8, and biology and frontier AI development tasks by Claude Opus 5. Anthropic says this "likely reduces" Opus 5.5's scores.
- Some margins come with error bars. On Terminal-Bench 4.0, the standard error is about 2.6 points for Opus 5.5. On Terminal-Bench-Science, it's 3.5 to 5 points per model.
- AutomationBench counted safeguard stops as failures. Zapier ran it without fallback models, which Anthropic says lowered Opus 5.5's score.
None of this is unusual. Every lab picks settings and sources that suit its launch. It does mean the chart is best read as Anthropic's best case, not a like-for-like test.
Which Leads in the Claude Opus 5.5 Benchmarks Survive the Noise?
Some are clear. Others are too close to call. Here's each row with the gap to the best rival shown.
| Benchmark | Opus 5.5 | Best rival shown | Gap | Our read |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1846 | Fable 5.1, 1735 | +111 | Clear lead |
| Terminal-Bench 4.0 | 66.4% | GPT-6 Astra, 57.9% | +8.5 | Clear on this setup, much smaller independently |
| CursorBench 4.0 | 57.8% | Fable 5.1, 51.8% | +6.0 | Clear, but no Astra number |
| Humanity's Last Exam | 67.7% | Fable 5.1, 65.6% | +2.1 | Small over Fable 5.1, large over Astra |
| FrontierCode v1.1 | 54.4% | GPT-6 Astra, 53.3% | +1.1 | Effectively a tie |
| OSWorld 2.0 | 81.8% | Fable 5.1, 80.7% | +1.1 | Tie, Anthropic models only |
| Chartography | 89.0% | Fable 5.1, 88.4% | +0.6 | Tie, Anthropic models only |
| AutomationBench | 40.0% | GPT-6 Astra, 41.4% | -1.4 | Effectively a tie |
| Terminal-Bench-Science 0.1 | 58.7% | GPT-6 Astra, 64.6% | -5.9 | Likely Astra lead, within wide error bars |
So the clear wins are knowledge work, Terminal-Bench on Anthropic's setup, and CursorBench. The rest are close, missing a key rival, or go to Astra.
Anthropic makes a similar point itself. Its announcement says that "benchmark margins have become a less reliable guide to real-world differences" at this level of capability.
How Do Independent Tests Compare?
They confirm the overall lead, with smaller gaps on some rows. Artificial Analysis runs every model on the same harness, which removes the mixed-source problem.
Here are the five models from the chart on its Intelligence Index, version 4.3.2, at each model's top setting. Figures come from Artificial Analysis's model pages for Opus 5.5, Opus 5 and GPT-5.6 Sol, plus its leaderboard.
| Model (top setting) | Intelligence Index | Terminal-Bench 4.0 | AutomationBench-AA | Cost per task |
|---|---|---|---|---|
| Claude Opus 5.5 | 58 | 60% | 70% | $5.98 |
| Claude Fable 5.1 | 53 | 52% | 59% | $7.63 |
| GPT-6 Astra | 53 | 59% | 68% | $3.26 |
| Claude Opus 5 | 51 | 49% | 57% | $5.86 |
| GPT-5.6 Sol | 47 | 40% | 60% | $1.99 |
Two differences from Anthropic's chart stand out:
- Terminal-Bench 4.0 is almost a tie. Anthropic's chart shows Opus 5.5 ahead of Astra by 8.5 points. In Artificial Analysis's run, the gap is 1 point, 60% to 59%. Anthropic's own models also score a few points lower in the independent run, which suggests a harness difference rather than anything hidden.
- AutomationBench flips. Zapier's run on Anthropic's chart has Astra ahead by 1.4 points. Artificial Analysis's version has Opus 5.5 ahead, 70% to 68%. Different versions and setups can move a close result either way.
The knowledge-work lead holds. GDPval-AA is Artificial Analysis's own test, so the chart and the independent figure match exactly at 1846.
Cost per task is worth noting too. At its top setting, Opus 5.5 costs more per task than GPT-6 Astra, because it writes far more tokens. Our Claude Opus 5.5 analysis covers cost by setting in detail.
Claude Opus 5.5 Benchmarks vs Opus 5: How Big Is the Upgrade?
Large, on every row. This is the comparison where the chart and independent data agree most clearly.
On Anthropic's chart, Opus 5.5 improves on Opus 5 by:
- 29.7 points on Terminal-Bench-Science, from 29.0% to 58.7%, the biggest jump on the chart.
- 14.1 points on Terminal-Bench 4.0, from 52.3% to 66.4%.
- 13.1 points on AutomationBench, from 26.9% to 40.0%.
- 11.2 points on CursorBench, from 46.6% to 57.8%.
- 138 rating points on GDPval-AA, from 1708 to 1846.
Independently, Artificial Analysis scores Opus 5.5 at 58 against 51 for Opus 5 at their top settings. At the high setting, Opus 5.5 scores 54 for $1.82 per task, while Opus 5 scores 48 for $3.61. That's a higher score at about half the cost.
For teams on Opus 5, the upgrade case is strong. We covered the older model in our Claude Opus 5 analysis.
What Does the Safeguard Footnote Mean in Production?
It means some of your requests may be answered by an older model. This is easy to miss in a benchmark footnote, but it matters for real systems.
Anthropic ran the benchmarks with production safeguards on. When they stepped in, cybersecurity tasks went to Claude Opus 4.8, and biology and frontier AI development tasks went to Claude Opus 5. The same safeguards apply to your API calls, unless you're approved through Anthropic's verification programs.
In practice:
- Security teams may find that many of their tasks aren't handled by Opus 5.5 at all.
- Life sciences teams face the same issue for biology work.
- Benchmark results for those areas partly reflect the fallback models, as Anthropic's footnote notes.
- Your logs should record which model actually answered, so you can tell when a fallback happened.
Artificial Analysis also labels its Opus 5.5 results "with fallback," so the independent scores include this behavior too.
Where Does GPT-5.6 Sol Fit?
It's the budget and speed option, not a contender for the top spot. GPT-5.6 Sol is an earlier OpenAI model, and it trails on every row of Anthropic's chart where it appears.
On the chart, Sol scores 37.3% on Terminal-Bench 4.0, 47.5% on FrontierCode, 41.7% on CursorBench and 22.4% on Terminal-Bench-Science. Its GDPval-AA rating of 1588 is still slightly ahead of GPT-6 Astra's 1542, which is a useful reminder that newer doesn't mean better on every task.
Independently, Artificial Analysis scores Sol at 47 at its top setting, for $1.99 per task. It's also fast, at about 82 output tokens per second, well ahead of the speeds Artificial Analysis has measured for Opus 5, Fable 5.1 and GPT-6 Astra.
So Sol still makes sense where speed matters more than peak quality, such as quick drafts, simple tool calls and high-volume jobs. For harder agent work, the newer models are well ahead. Note that Opus 5.5 at medium scored higher than Sol's top setting on Artificial Analysis's index, at a lower cost per task.
Which Model Should You Pick for Each Task?
Use the rows that match your work, and weight the independent data more. Here's a starting map:
- Knowledge-work agents: Opus 5.5. It leads on both the chart and independent tests by a wide margin.
- Terminal and coding agents: Opus 5.5 or GPT-6 Astra. The chart favors Opus 5.5, but independent results are close to a tie.
- Business workflow automation: a toss-up between Opus 5.5 and GPT-6 Astra. The two sources disagree on the leader.
- Science research agents: GPT-6 Astra, based on Anthropic's own chart, though the error bars are wide.
- Computer use and chart reading: Opus 5.5 leads Anthropic's other models, but there's no cross-vendor data.
- Cost-sensitive work: Opus 5.5 at medium scored 51 for $1.34 per task on Artificial Analysis's index. GPT-5.6 Sol at its top setting scored 47 for $1.99.
For a wider view of how these models trade off on price, see our Grok 4.7 vs Fable 5.1 vs GPT-6 Astra comparison and our earlier look at Kimi K3 vs Opus 5 vs GPT-5.6 Sol.
How Should You Test Claude Opus 5.5 Benchmarks Against Your Own Work?
Treat the chart as a shortlist, then measure on your tasks. Here's a short plan:
- Pick the chart rows closest to your work, and build 50 to 200 real tasks that match them.
- Run Opus 5.5 at medium and high, not just max, since most teams won't pay for max on every call.
- Run your current model and GPT-6 Astra on the same tasks, with the same prompts and tools.
- Record success, cost and time per task, and which model actually answered each one.
- Check the close rows twice, since a 1-point lead on a benchmark rarely shows up in real work.
Our guide to AI agent evaluation covers how to build and maintain a test set like this.
How Van Data Team Helps Teams Check Claude Opus 5.5 Benchmarks Against Real Work
We help teams turn vendor benchmark charts into decisions based on their own data. That means building task sets from real work, running models side by side at several settings, and measuring success rate, cost and time per finished task.
Launch charts, including the Claude Opus 5.5 benchmarks, are a good place to start and a poor place to stop. If you want a model stack that holds up after the launch week buzz, our work on AI agent evaluation and AI agent development cost is a good place to start.
Article FAQ
Questions readers usually ask next.
These short answers clarify the practical follow-up questions that often come after the main article.
Need a similar system?
If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.
Book your free workflow review here.
