Skip to main content
Back to insights

September 29, 2026

Claude Sonnet 5.5 vs GPT-6 Luna: Cost, Speed and Accuracy

Claude Sonnet 5.5 vs GPT-6 Luna: Luna costs 20x less per token, Sonnet is far stronger at the top. Independent tests show which model to use for each job.

By Tran Tien Van9 min read

Article focus

GPT-6 Luna costs 20 times less per token than Claude Sonnet 5.5. At their top settings, Sonnet 5.5 is far more capable. At similar scores, Luna is about six times cheaper per task, but Sonnet 5.5 answers faster and makes far fewer factual mistakes. Here's what independent tests show, and which model fits which job.

Claude Sonnet 5.5 vs GPT-6 Luna isn't a fair fight on price: Luna costs 20 times less per token. At their top settings, Sonnet 5.5 is far more capable, scoring 56 to Luna's 37 on an independent index. At similar scores, Luna is about six times cheaper per task and wins most tests, but Sonnet 5.5 answers faster and makes far fewer factual mistakes. Pick by the job.

Key Takeaways

  • GPT-6 Luna costs $0.10 input and $0.50 output per million tokens. Claude Sonnet 5.5 costs $2 and $10, which is 20 times more.
  • At their top settings, Sonnet 5.5 scores 56 on the Artificial Analysis index and Luna scores 37. Sonnet 5.5 leads on nine of ten tests.
  • At similar scores, Sonnet 5.5 at low effort against Luna at max, Luna wins eight of ten tests at about a sixth of the cost.
  • Sonnet 5.5 makes far fewer factual errors, and at low effort it starts answering in under a second. Luna at max takes about 94 seconds.
  • Most teams should route: Luna for bulk work with the facts supplied, Sonnet 5.5 for coding, factual answers and hard tasks.

What Are Claude Sonnet 5.5 and GPT-6 Luna?

Two new models from different price tiers, released a week apart.

GPT-6 Luna is OpenAI's cheapest GPT-6 model, released on September 22, 2026. It replaced GPT-5.6 Luna at about half the price, as we covered in our GPT-6 Sol and Luna guide. Claude Sonnet 5.5 is Anthropic's mid-tier model, released on September 28, 2026. We covered it in detail in our Claude Sonnet 5.5 analysis.

Claude Sonnet 5.5GPT-6 Luna
ReleasedSeptember 28, 2026September 22, 2026
Input price per million tokens$2.00$0.10
Output price per million tokens$10.00$0.50
Cached input per million tokens$0.20$0.01
Context window1 million tokens1 million tokens
Top Artificial Analysis index score5637

Prices and context windows come from Artificial Analysis's comparison page. Anthropic lists the same prices on its launch page.

So why compare them at all? Because teams often choose between "the cheap model" and "the good model" for the same job. Anthropic's closer rival to Luna is likely Haiku 5.5, which VentureBeat reports is due in the coming weeks. Until then, many teams are weighing exactly this pair.

Claude Sonnet 5.5 vs GPT-6 Luna at Their Top Settings

Sonnet 5.5 wins almost everything, and costs about 100 times more per task to do it.

Independent data: Here are both models at max effort on the Artificial Analysis Intelligence Index, version 4.3.2:

TestSonnet 5.5 (max)Luna (max)
Intelligence Index5637
GDPval-AA (knowledge work)18441367
AutomationBench-AA71%53%
Terminal-Bench 4.064%13%
SciCode61%55%
Humanity's Last Exam55%39%
CritPt (research physics)31%19%
AA-Omniscience (facts)321
AA-LCR (long context)83%83%
Cost per task$7.60$0.07

The biggest gaps are in terminal work, at 64% to 13%, and factual recall, at 32 to 1. The one tie is long-context reasoning, at 83% each. That tie matters: when the answer is inside a long document you provide, Luna does as well as Sonnet 5.5 here.

Token use explains much of the gap. At max effort, Sonnet 5.5 used about 193,000 output tokens per task against Luna's 51,000, so running the full index cost $8,977 for Sonnet 5.5 and just $122 for Luna.

Claude Sonnet 5.5 vs GPT-6 Luna at Similar Scores

This is the fairer fight. Sonnet 5.5 at low effort scores 36, and Luna at max scores 37.

Independent data: Here's the head-to-head at those settings:

TestSonnet 5.5 (low)Luna (max)
Intelligence Index3637
GDPval-AA11681367
AutomationBench-AA50%53%
Terminal-Bench 4.021%13%
SciCode49%55%
AA-Omniscience191
AA-LCR76%83%
Cost per task$0.41$0.07
Time to first token0.92 seconds94 seconds

Across all ten tests on the index, Luna wins eight. That includes knowledge work, science coding, long documents and the hardest exam questions. It does that for about a sixth of the cost.

Sonnet 5.5 wins the other two. Both matter. It does better on terminal tasks, at 21% to 13%, and it's far more reliable on facts, at 19 to 1.

It also starts answering in under a second, while Luna at max thinks for about a minute and a half first.

Our view: for batch jobs, Luna looks like the better buy at this level. For anything live, factual or terminal-based, the extra cost of Sonnet 5.5 buys something real: fewer wrong answers, and replies that start right away.

Why Does the Factual Accuracy Gap Matter?

Because a cheap wrong answer can cost more than an expensive right one.

AA-Omniscience asks factual questions and rewards correct answers. It subtracts points for wrong ones, and a model that says "I don't know" loses nothing. A score near zero means the model gives about as many wrong answers as right ones when it answers from memory.

Luna scores 1 at max and minus 6 at high. Sonnet 5.5 scores 19 at low and 32 at max. So when Luna answers a factual question from memory, it's close to a coin flip on whether it's right.

That doesn't make Luna a bad model. It makes it a model to use with sources. Its long-context score of 83% shows it reads well. So give it the documents, the database rows or the search results, and ask it to answer from those.

  • Good fit for Luna: "Summarize this contract," "extract the totals from these invoices," "classify these support tickets."
  • Poor fit for Luna: "What's the capital gains rule in this country?" or "Which drug interacts with this one?" asked with no source attached.

Claude Sonnet 5.5 vs GPT-6 Luna on Speed

Luna writes faster, but Sonnet 5.5 can start sooner, and which one matters depends on what your product does.

  • Output speed. Luna runs at about 141 to 154 tokens per second across settings, against about 85 to 139 for Sonnet 5.5.
  • Time to first token at low settings. Sonnet 5.5 at low effort starts in about 0.92 seconds, while Luna at high effort takes about 5.8 seconds, roughly six times longer.
  • Time to first token at max. Luna takes about 94 seconds. Sonnet 5.5 takes about 328 seconds, because both models think before answering.

For a chat window, the first token is what users feel. A reply that starts in one second feels instant, while one that starts in six seconds feels slow. For a batch job that runs overnight, only total throughput and cost matter, and Luna wins there.

How Much Would Each Cost at Scale?

Here's a simple way to picture the gap, using Artificial Analysis's cost per task. These are hard test tasks, so your real tasks may cost much less. The ratios are what matter.

Model and settingCost per taskCost for 10,000 tasks
GPT-6 Luna, high$0.03$300
GPT-6 Luna, max$0.07$700
Claude Sonnet 5.5, low$0.41$4,100
Claude Sonnet 5.5, high$1.08$10,800
Claude Sonnet 5.5, max$7.60$76,000

The Luna figures come from its results by setting. Caching widens the gap further for repeated prompts, since Luna's cached input costs $0.01 per million tokens against $0.20 for Sonnet 5.5. Our guide to token efficiency explains why tokens per task, not price per token, drives the bill.

The honest read: for millions of simple tasks, Luna's price is hard to argue with. For a few thousand hard ones, the cost of Sonnet 5.5 is often small next to the cost of a wrong answer.

What Are the Main Risks With Each Model?

Every model has failure modes. Here are the ones the tests point to, so you can plan around them.

GPT-6 Luna:

  • Wrong facts from memory. Its AA-Omniscience score is close to zero at max and below zero at high, so don't let it answer factual questions without a source attached.
  • Long waits at max effort. About 94 seconds before the first token rules out max effort for anything interactive.
  • Weak terminal and coding work. At 13% on Terminal-Bench 4.0 at max, it isn't the model for coding agents.
  • Hidden cost from retries. If a cheap answer fails a check and you retry on a bigger model, you pay twice. Count retries in your cost per task.

Claude Sonnet 5.5:

  • Cost at max effort. At $7.60 per task on the index, max effort costs more than Opus 5.5's $5.98, as our Sonnet 5.5 analysis shows.
  • Heavy token use at high settings. About 193,000 output tokens per task at max adds up fast on long jobs.
  • Cyber fallback. Anthropic says higher-risk security tasks visibly fall back to Sonnet 5, which changes quality mid-task.
  • Slower writing. Its output speed trails Luna's at every setting, which matters for long outputs.

None of these rule a model out; they tell you which routes need a check, a fallback or a lower setting before you ship.

Which Should You Pick: Claude Sonnet 5.5 vs GPT-6 Luna?

Match the model to the job. Here's how we'd route common work, based on the tests above:

JobBetter pickWhy
Extracting fields from documentsGPT-6 LunaReads long input well, at a fraction of the cost
Classifying tickets or emailsGPT-6 LunaHigh volume, and the text holds the answer
Summarizing long reportsGPT-6 LunaTies Sonnet 5.5 on long-context reasoning
Coding and terminal agentsClaude Sonnet 5.5Far ahead on Terminal-Bench 4.0 at every setting
Answering from the model's own knowledgeClaude Sonnet 5.5Much more reliable on facts
Live chat with a fast first replyClaude Sonnet 5.5, lowStarts answering in under a second
Hard analysis and knowledge workClaude Sonnet 5.5, high or aboveMuch higher scores on GDPval-AA

For the hardest work, it may be worth looking one tier up. Our GPT-6 Sol and Luna vs Claude Opus 5.5 comparison covers that matchup.

Why Not Use Both?

Many teams should. A two-model setup lets each model do what it's good at.

  • Send bulk work to Luna. Extraction, tagging and summaries with the source attached go to the cheap model, which is usually where most of the volume sits.
  • Send the rest to Sonnet 5.5. Coding, open questions, anything customer-facing that must be right.
  • Escalate on doubt. If Luna's output fails a check, such as a missing field or a low confidence score, retry the task on Sonnet 5.5.
  • Keep the prompts close. Write prompts that work on both, so you can switch a route without rewriting it.

The savings can be large. If most of your volume is simple and a small share needs the stronger model, routing can cut your bill sharply without hurting quality. Test it on your own traffic first.

How Should You Test Before Choosing?

With your own tasks, not ours. A short plan:

  • Build a test set of 50 to 200 real tasks with known good answers. Our guide to AI agent evaluation shows how.
  • Run both models at two or three settings each, such as Luna at high and max, and Sonnet 5.5 at low and high.
  • Score accuracy, cost per task and time to first token for each run.
  • Check factual answers by hand on a sample, since that's where Luna is weakest.
  • Decide per route, not for the whole product.

How Van Data Team Helps Teams Pick and Route Models

We help teams choose models based on their own data. That means test sets from real tasks, cost and latency tracking per route, and routing rules that send each job to the cheapest model that does it well.

Claude Sonnet 5.5 vs GPT-6 Luna is a good example of why one model rarely fits every job. If you want help setting up that routing, our AI governance guide is a good place to start.

Article FAQ

Questions readers usually ask next.

These short answers clarify the practical follow-up questions that often come after the main article.

Need a similar system?

If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.

Book your free workflow review here.