Skip to main content
Back to insights

September 24, 2026

GDPval-AA Chart Fact-Check: Opus 5.5 vs GPT-6 Astra on Cost

A viral GDPval-AA chart says Opus 5.5 crushes GPT-6 Astra at a tenth of the cost. We check every claim, fix the numbers, and show the fairer comparison.

By Tran Tien Van9 min read

Article focus

A viral post says Claude Opus 5.5 at medium effort crushes GPT-6 Astra at max on GDPval-AA, at a tenth of the cost. The direction is right, but several numbers are off, one statistic comes from the wrong source, and the post compares the wrong points. Here's the fact-check, and the fairer comparison.

A viral GDPval-AA chart claims Claude Opus 5.5 at medium effort crushes GPT-6 Astra at max for a tenth of the cost. The direction is right: Opus 5.5 scores higher at every cost where both appear. But medium beats Astra's max by only 34 Elo points, at about a fifth of the cost, not a tenth. The fairer comparison, at matched cost, shows a much bigger gap.

Key Takeaways

  • On Artificial Analysis's GDPval-AA test, Claude Opus 5.5 scores 1846 Elo at max effort. GPT-6 Astra scores 1542 at max.
  • The viral post's main point holds: on Anthropic's chart, Opus 5.5 is higher than Astra at every cost level where both appear.
  • But Opus 5.5 at medium beats Astra at max by just 34 points, at about a fifth of the cost. The post says a tenth.
  • The post's Terminal-Bench figures come from Anthropic's own chart, not Artificial Analysis. Independently, the two models are close to tied there.
  • At matched cost, the GDPval-AA gap is 210 to 280 points, a stronger finding than the one the post chose.

What Does the GDPval-AA Chart Show?

It plots each model's score against its cost per task, at each effort setting. The chart comes from Anthropic's Claude Opus 5.5 announcement, in its knowledge work section.

The vertical axis is Elo, a rating from head-to-head comparisons. The horizontal axis is estimated cost per task in US dollars, on a log scale, so each step to the right multiplies the cost. Anthropic plots Opus 5.5 in orange and GPT-6 Astra in gray, with Fable 5.1, Opus 5 and GPT-5.6 Sol available but grayed out.

Here are the points, as read from Anthropic's chart and transcribed by Kingy AI. The top two Opus 5.5 and Astra scores match Artificial Analysis's GDPval-AA leaderboard exactly.

Effort settingOpus 5.5 cost per taskOpus 5.5 EloGPT-6 Astra cost per taskGPT-6 Astra Elo
LowAbout $0.211224About $0.851366
MediumAbout $0.861576About $1.821468
HighAbout $1.541692About $2.431485
XhighAbout $4.211820About $3.041516
MaxAbout $8.921846About $4.531542

Anthropic's own caption says that at its default medium setting, "Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task." Keep that phrase in mind. The viral post goes further than Anthropic does.

What Is GDPval-AA, and How Is It Scored?

It's a test of real-world work, not trivia. Artificial Analysis describes the method on its leaderboard page:

  • The tasks: 220 tasks from GDPval, a set OpenAI developed with industry professionals, covering 44 occupations across 9 industries.
  • The work: models run in an agent loop with shell access and web browsing, and produce documents, slides, diagrams and spreadsheets.
  • The scoring: two models' outputs are anonymized, and an AI judge picks the better one. Those results become an Elo rating per model.
  • The anchor: ratings are pinned to DeepSeek V4.1 Flash at max, set at 1,600.
  • The error margins: about 15 to 26 points either way for the top models.

Two details matter. First, the tasks come from OpenAI's own benchmark, so this isn't a test built to favor Anthropic. Second, the leaderboard page doesn't name the judge model. AI judges can have their own biases, so a human spot-check is worth doing for any high-stakes use.

How Big Is a 34-Point GDPval-AA Gap?

Small. On a standard Elo scale, a rating gap tells you how often the higher-rated model would win a head-to-head comparison. Artificial Analysis doesn't state its exact scale, so treat these as rough guides:

  • 34 points, Opus 5.5 at medium against Astra at max: about a 55% win rate, close to a coin flip.
  • 210 points, Opus 5.5 at medium against Astra at low, at the same cost: about 77%.
  • 278 points, Opus 5.5 at xhigh against Astra at max, at similar cost: about 83%.
  • 304 points, both models at max: about 85%.

The error margins matter too. Astra at max is 1542, give or take 25 points. A 34-point lead is not far outside that range. Calling it a crushing win overstates it.

Which Claims in the Viral GDPval-AA Post Hold Up?

Most of the numbers are close, but a few are wrong, and one comes from the wrong source. Here's each claim checked:

Claim in the postWhat we foundVerdict
Opus 5.5 low: about $0.20 for 1220 EloAbout $0.21 for 1224Correct
Medium: about $0.80 for 1580. High: 1690About $0.86 for 1576. High: 1692Correct
Xhigh: $5 for 1820. Max: $10 for 1850About $4.21 and $8.92. Elo 1820 and 1846Roughly right, costs rounded up
Astra only plotted from about $1, peak 1540Starts at about $0.85, peak 1542Roughly right
Opus 5.5 medium crushes Astra max at about a tenth of the cost34 points higher, at about a fifth of the costOverstated
Artificial Analysis: GDPval-AA 1846 vs 1542, AA-Briefcase 1822 vs 1569Matches Artificial AnalysisCorrect
Artificial Analysis: Terminal-Bench 4.0 is 66.4% vs 57.9%Those are Anthropic's figures. Artificial Analysis shows 60% vs 59%Wrong source
Released September 22 at $4 and $20 per million tokensMatches Anthropic's announcementCorrect
About 40% cheaper than Opus 5Per-token prices are 20% lower. The 40% is Anthropic's claim for typical workloadsMixed up
Anthropic wins at both ends, cheapest and strongestTrue against Astra on this chart, which leaves out several rivalsTrue, but narrow

The Terminal-Bench mix-up matters most. Anthropic's chart shows Opus 5.5 at xhigh against Astra's figure as reported by OpenAI, and the gap is 8.5 points. In Artificial Analysis's own run, with one harness for both, the gap is 1 point. We covered that difference in our Claude Opus 5.5 benchmarks breakdown.

What Is the Fairer GDPval-AA Comparison?

Compare at the same cost, not across settings. The post compared Opus 5.5 at medium with Astra at max, which mixes a cheap setting with an expensive one. Matching costs tells a clearer story:

  • At about $0.85 per task: Opus 5.5 at medium scores 1576. Astra at low scores 1366. That's 210 points.
  • At about $1.50 to $1.80: Opus 5.5 at high scores 1692. Astra at medium scores 1468. That's 224 points.
  • At about $4.20 to $4.50: Opus 5.5 at xhigh scores 1820. Astra at max scores 1542. That's 278 points.

So the post undersold the real finding while overstating its numbers. On this test, the gap at matched cost is large and consistent. That's a stronger and more defensible claim than "medium beats max."

One caution: on Anthropic's chart, Opus 5.5's costs are estimates. Artificial Analysis says its costs are built from input, cached, reasoning and answer token counts where available. Your own costs will depend on your prompts, caching and how much each model writes.

What Does the GDPval-AA Chart Leave Out?

Quite a lot. The chart is a fair picture of one test, but it's still a vendor's chart.

  • Other rivals. It plots only Anthropic and OpenAI models. On the same leaderboard, Grok 4.7 scores 1695, well above Astra. GPT-6 Sol scores 1487 and GPT-6 Luna 1367.
  • Other tests. On Artificial Analysis's other tests, Astra comes close or leads. It leads on reasoning over long professional documents, and it's within a point or two on Terminal-Bench and business app automation.
  • Speed. Artificial Analysis hadn't published Opus 5.5's speed when we checked. It also writes far more tokens per task than Astra at the top setting.
  • Safeguards. Anthropic routes most cybersecurity tasks to an older model and restricts advanced biology work, which matters for some teams.

The post's line that Astra "isn't weak, it's just being left behind" fits GDPval-AA well. Across all ten of Artificial Analysis's tests, the picture is closer. Our GPT-6 Sol vs Claude Opus 5.5 comparison shows how cheaper models fit in.

Where Would Grok 4.7 and GPT-6 Sol Sit on This Chart?

Closer to the top than the chart suggests, in Grok's case. Artificial Analysis rates Grok 4.7 at 1695 at xhigh and 1694 at high. That's level with Opus 5.5 at high, which scores 1692 on Anthropic's chart, and well above GPT-6 Astra at any setting.

OpenAI's cheaper models sit lower. GPT-6 Sol scores 1487 at max, just below Astra's max, and GPT-6 Luna scores 1367. Fable 5.1, grayed out on Anthropic's chart, scores 1735.

Anthropic's chart doesn't show their cost on this test, so we can't place them exactly. On Artificial Analysis's overall index, Grok 4.7 at high costs about $2.73 per task and Opus 5.5 at high about $1.82. Those are averages across ten tests, though, not GDPval-AA alone.

The lesson is simple. A chart with two vendors can't tell you who leads the whole market. For knowledge work, Grok 4.7 is a real option this chart doesn't show. Our Grok 4.7 analysis covers its price and token use.

Is Opus 5.5 Worth Upgrading to in Claude Code?

The post ends by asking this. The honest answer depends on your work:

  • For documents, analysis and knowledge work: the GDPval-AA evidence is strong. Opus 5.5 leads clearly at every cost level.
  • For coding: the independent evidence is closer. On Artificial Analysis's Terminal-Bench run, Opus 5.5 and Astra are nearly tied. In one hands-on test by DataCamp, GPT-6 Sol beat Opus 5.5 on quality at a third of the cost.
  • For cost: medium and high settings give the best value. Max adds little for a lot more money, at about 26 points for more than twice the cost of xhigh.
  • Against Opus 5: the upgrade case is strong on both price and scores, as we covered in our Claude Opus 5.5 analysis.

The best test is your own repository. Run your usual tasks at medium and high, and compare against your current model on success rate, cost and time.

How Should You Read Elo-vs-Cost Charts?

With a few checks. These charts are useful, but easy to misread.

  • Mind the log scale. Equal steps on the axis mean multiplying the cost, not adding to it. Points that look close can differ by several times.
  • Compare at matched cost. Draw a vertical line and see which model is higher at the same price.
  • Check the error margins. Gaps smaller than about two error margins may not be real.
  • Ask who made it. A vendor chooses which models and settings appear.
  • Check what's missing. Rivals, other tests, speed and limits can all change the decision.
  • Check if costs are estimated. Your real costs depend on your own prompts and token use.

How Can You Run a GDPval-Style Test on Your Own Work?

Use the same idea: blind, head-to-head comparisons on real tasks.

  • Pick 30 to 100 real work tasks, such as reports, analyses or slide outlines your team actually produces.
  • Run two models on each task, at the settings you'd use in production.
  • Hide which model wrote which answer, and have a reviewer or an AI judge pick the better one.
  • Spot-check the judge with a human on a sample, since AI judges can be biased.
  • Record the cost of each answer, and compare win rate against cost.

Our guide to AI agent evaluation covers how to build and maintain a test set like this.

How Van Data Team Helps Teams Test Models on Real Work

We help teams go beyond vendor charts with tests built from their own work. That means blind head-to-head comparisons, human spot-checks of AI judges, and cost per task measured at the settings you'd actually use.

A viral chart can start a good conversation. Your own GDPval-AA-style test should end it. Our work on AI agent evaluation and AI agent development cost is a good place to start.

Article FAQ

Questions readers usually ask next.

These short answers clarify the practical follow-up questions that often come after the main article.

Need a similar system?

If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.

Book your free workflow review here.