Skip to main content
Back to insights

September 24, 2026

GPT-6 Sol vs Claude Opus 5.5, Plus Luna: Best Value per Task

GPT-6 Sol vs Claude Opus 5.5, plus Luna, at matched budgets: who scores higher per dollar, a hands-on coding test, speed, and how to route between them.

By Tran Tien Van9 min read

Article focus

Claude Opus 5.5 scores higher than GPT-6 Sol and Luna on every independent test at their top settings. But at the same budget per task, Sol wins up to a mid-level score, and Luna has no rival at its price. Here's the comparison at matched cost, a hands-on test, and a routing plan that uses all three.

GPT-6 Sol vs Claude Opus 5.5 comes down to budget. On independent tests, Opus 5.5 scores higher at its top setting, 58 against 48, and leads all ten underlying tests. But at the same cost per task, GPT-6 Sol wins up to a mid-level score, and GPT-6 Luna has no Anthropic rival at its price. The practical answer is to route work across all three.

Key Takeaways

  • On the independent Artificial Analysis index, Claude Opus 5.5 scores 58 at its top setting, GPT-6 Sol 48 and GPT-6 Luna 37.
  • Per token, Sol costs half as much as Opus 5.5, and Luna costs a fortieth as much.
  • At matched cost, Sol wins up to a score of about 44. Above 48, only Opus 5.5 gets there.
  • In one hands-on coding test at the high setting, Sol cost $0.26 and scored higher on quality than Opus 5.5 at $0.87.
  • Anthropic's model at Sol's price, Claude Sonnet 5, scores far lower per dollar. A three-model routing plan beats any single choice.

How Do GPT-6 Sol, Luna and Claude Opus 5.5 Compare at a Glance?

Here are the three models side by side at their top settings. Prices come from each vendor. Scores, speed and cost per task come from Artificial Analysis's pages for Opus 5.5, GPT-6 Sol and GPT-6 Luna, checked on September 24, 2026.

FeatureClaude Opus 5.5GPT-6 SolGPT-6 Luna
Price per million tokens (input / output)$4 / $20$2 / $10$0.10 / $0.50
Cached input per million tokens$0.20$0.20$0.01
Context window1 million tokensAbout 872,000 tokens1 million tokens
Intelligence Index (top setting)584837
Cost per task (top setting)$5.98$1.06$0.07
Output tokens per task (top setting)About 119,000About 31,000About 51,000
Output speedNot yet publishedAbout 125 tokens/sAbout 154 tokens/s

All three launched on September 22, 2026, within about 90 minutes of each other. We covered each launch separately in our Claude Opus 5.5 analysis and our GPT-6 Sol and Luna analysis.

GPT-6 Sol vs Claude Opus 5.5: Which Wins at the Same Budget?

Sol wins at low and mid budgets. Opus 5.5 wins once you need a high score. Comparing top settings alone hides this, so here's every setting sorted by cost per task on Artificial Analysis's index:

Model and settingIntelligence IndexCost per task
GPT-6 Luna (high)32$0.03
GPT-6 Luna (max)37$0.07
GPT-6 Sol (low)34$0.13
GPT-6 Sol (medium)40$0.25
GPT-6 Sol (high)43$0.37
GPT-6 Sol (xhigh)44$0.53
Claude Opus 5.5 (low)42$0.55
GPT-6 Sol (max)48$1.06
Claude Opus 5.5 (medium)51$1.34
Claude Opus 5.5 (high)54$1.82
Claude Opus 5.5 (xhigh)56$3.46
Claude Opus 5.5 (max)58$5.98

Here's how to read it:

  • Up to a score of about 44, Sol is cheaper. Sol at xhigh scores 44 for $0.53, beating Opus 5.5 at low, which scores 42 for $0.55.
  • Every Sol setting below max costs less than Opus 5.5's cheapest setting.
  • Around $1 to $1.35 per task, it's close. Sol at max scores 48 for $1.06. Opus 5.5 at medium scores 51 for $1.34.
  • Above 48, only Opus 5.5 gets there. Its high setting scores 54 for $1.82, which Sol can't reach at any price.

So the question isn't which model is better. It's what score your task actually needs.

GPT-6 Sol vs Claude Opus 5.5: Which Scores Higher, Test by Test?

Opus 5.5, on all ten tests, though some margins are tiny. Here's Artificial Analysis's breakdown with each model at its top setting, from its Opus 5.5 vs Fable 5.1 and Luna vs Sol comparisons:

Test (Artificial Analysis)Opus 5.5GPT-6 SolGPT-6 Luna
GDPval-AA v2.1 (agent tasks from 44 occupations)184614871367
AA-Briefcase v1.1 (multi-week knowledge work)182214831299
AutomationBench-AA (business app workflows)70%62%53%
Terminal-Bench 4.0 (terminal tasks)60%44%13%
SciCode (scientific Python)67%58%55%
Humanity's Last Exam (hard questions)61%48%39%
AA-Omniscience (factual reliability)46271
CritPt (research-level physics)32%31%19%
GDP.pdf (long professional documents)26%25%20%
AA-LCR v1.1 (long-context reasoning)85%84%83%

The big gaps are in knowledge work, terminal tasks and hard questions. On research physics, long documents and long-context reasoning, Opus 5.5 and Sol are within a point of each other. Luna also comes close on long-context reasoning, at a tiny fraction of the cost.

Luna's factual-reliability score of 1 stands out. That test penalizes wrong answers, so Luna is best used with the facts supplied, through retrieval or documents.

What Happened in a Hands-On Coding Test?

Sol came out ahead in one independent test. DataCamp asked both models to build a playable Tetris-style game in a single HTML file, with a twist: gravity flips every 20 seconds. Both ran once at the high setting with the same tools and prompt.

  • Cost: Sol spent $0.26. Opus 5.5 spent $0.87.
  • Quality: Sol averaged 5.0 out of 5 on DataCamp's scoring. Opus 5.5 averaged 4.3.
  • Approach: Opus 5.5 finished in 3 turns with 2 tool calls, thinking harder per step. Sol took 6 turns and 9 tool calls, but each was cheap.
  • Both got the hard part right. DataCamp gave both full marks on the gravity-flip mechanics.

DataCamp's summary is that Opus 5.5 "thought roughly three times as hard per turn, reached a correct build in half the turns, and stopped," while Sol "iterated in the file instead, cheaply, and came out slightly ahead on feel."

It's one run of one task at one setting, as DataCamp itself notes. But it's a useful reminder: for contained coding jobs, a cheaper model that iterates can match or beat a stronger one that stops early.

Where Does GPT-6 Luna Fit Against Claude?

Below everything Anthropic currently offers on price. Luna's fair comparison isn't Opus 5.5, but Anthropic's smaller models.

  • Claude Sonnet 5 lists at $2 input and $10 output per million tokens, the same as GPT-6 Sol. On Artificial Analysis's index, it scores 38 at its top setting for $5.09 per task, and 32 at high for $1.79.
  • GPT-6 Luna scores 37 at its top setting for $0.07 per task, about the same score as Sonnet 5 for a small fraction of the cost.
  • GPT-6 Sol beats Sonnet 5 at every level. Sol at low scores 34 for $0.13, a score Sonnet 5 only reaches at xhigh, for $2.87.
  • Claude Haiku 4.5, Anthropic's cheapest current model, lists at $1 input and $5 output per million tokens, ten times Luna's per-token price.

That's a clear win for OpenAI in the cheap and mid tiers right now. Sonnet 5 is very verbose on these tests, which drives up its cost per task. We covered its launch in our Claude Sonnet 5 analysis.

GPT-6 Sol vs Claude Opus 5.5: How Do Speed and Tokens Compare?

Sol and Luna are faster and write less. Artificial Analysis measures Sol at about 125 tokens per second and Luna at about 154. It hadn't published Opus 5.5's speed when we checked.

Token use differs even more. At their top settings, Opus 5.5 writes about 119,000 output tokens per task, Sol about 31,000 and Luna about 51,000. In DataCamp's test, Opus 5.5 used about 22,500 reasoning tokens against Sol's 8,100.

That matters for waiting time as well as cost. A model that writes four times as many tokens takes longer to finish, even at a similar speed. For interactive tools, Sol and Luna will usually feel quicker. Simon Willison also found that Opus 5.5 at max "over-thinks to breaking point" on one test, running out of output room.

What About Vendor Claims and Safeguards?

Neither launch compared these models directly, and their own tables can't be lined up fairly.

  • Anthropic's chart compared Opus 5.5 with the older GPT-5.6 Sol, not GPT-6 Sol, which launched after it. We went through that chart in our Claude Opus 5.5 benchmarks breakdown.
  • OpenAI said Sol and Luna "substantially" outperform Anthropic's Fable and Opus models on various tasks, TechCrunch reported. Opus 5.5 came out 90 minutes earlier, so that claim likely didn't include it.
  • Cross-vendor numbers don't match up. Both companies report OSWorld and AutomationBench, but with different versions and settings, so the figures shouldn't be compared side by side.

Safeguards also differ. Anthropic says most cybersecurity tasks sent to Opus 5.5 are handled by the older Opus 4.8, and advanced biology work by Opus 5, unless you're approved through its verification programs. We didn't find a similar handoff described for Sol or Luna. If your work touches those areas, test which model actually answers.

How Should You Route Work Across Luna, Sol and Opus 5.5?

Use all three, each for what it does best. A simple routing plan looks like this:

  • Luna first. Classify each request, and let Luna handle summaries, extraction and simple questions, with facts supplied.
  • Sol for the middle. Send everyday coding help, tool calls and routine agent steps to Sol at medium or high.
  • Opus 5.5 for the hardest work. Send long, complex agent tasks, deep reasoning and knowledge work to Opus 5.5 at medium or high.
  • Escalate on failure. If a cheaper model fails a check, retry the task one level up.

The savings can be large. Say 70% of tasks go to Luna at high, 20% to Sol at high and 10% to Opus 5.5 at high. Using Artificial Analysis's cost per task, that averages about $0.28 per task, against $1.82 if everything went to Opus 5.5. That's about 85% less, provided quality holds on your tasks.

We covered the same routing idea for even cheaper classifiers in our guide to fast System One models for agents.

GPT-6 Sol vs Claude Opus 5.5: Which Should You Choose?

If you can only pick one, match it to your hardest common task:

  • Most work is routine or mid-level: GPT-6 Sol. It reaches good scores for much less, and it's fast.
  • Most work is complex agent tasks or knowledge work: Claude Opus 5.5 at medium or high. It's the only one of the three that reaches top scores.
  • Most work is high-volume summaries or extraction: GPT-6 Luna, with retrieval.
  • You're on Claude Sonnet 5 for cost reasons: test GPT-6 Sol, which scores higher for less on independent tests.
  • You're in security or biology: check how Anthropic's safeguards affect your tasks before choosing Opus 5.5.

How Should You Test GPT-6 Sol vs Claude Opus 5.5 on Your Work?

Measure cost per successful task at several settings, not the leaderboard.

  • Pick 50 to 200 real tasks with known good answers, split by difficulty.
  • Run Luna, Sol and Opus 5.5 at medium and high, with the same prompts and tools.
  • Record success, cost, tokens and wall-clock time for every task.
  • Find the cheapest model and setting that clears your quality bar for each task type.
  • Log which model answered, so safeguard handoffs don't hide in your results.

Our guide to AI agent evaluation covers how to build and maintain a test set like this.

How Van Data Team Helps Teams Route Between OpenAI and Anthropic Models

We help teams build model routing that sends each task to the cheapest model that does it well, across vendors. That means building task sets from real work, measuring cost and quality per task, and setting up escalation when a cheaper model falls short.

The GPT-6 Sol vs Claude Opus 5.5 choice doesn't have to be either-or. If you want lower costs without losing quality on the hard cases, our work on AI agent evaluation and AI agent development cost is a good place to start.

Article FAQ

Questions readers usually ask next.

These short answers clarify the practical follow-up questions that often come after the main article.

Need a similar system?

If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.

Book your free workflow review here.