October 1, 2026
Gemini 4 Argon Benchmarks: What Independent Tests Show
Gemini 4 Argon tops Google's own chart, but independent tests tie it with GPT-6 Astra and put Opus 5.5 ahead. What's real, what it costs and who can use it.
Article focus
Google announced Gemini 4 Argon on September 30, and its launch chart puts it ahead of GPT-6 Astra and Claude Opus 5.5 on most tests. Independent results are more mixed: a tie with Astra, a clear lead on factual reliability, and a cost edge that depends on introductory pricing. Here's what we know, and what builders should do while access is limited.
Section guide
Gemini 4 Argon is Google's new frontier model, announced on September 30, 2026, and its launch chart puts it ahead of GPT-6 Astra and Claude Opus 5.5 on most tests. Independent tests are more mixed. Argon ties Astra on the Artificial Analysis index and trails Opus 5.5, but it hallucinates far less than its rivals. Most people can't use it yet.
Key Takeaways
- Google announced Gemini 4 Argon on September 30, 2026. Only trusted cyber defenders in its Fairwind Program can use it so far, with paid API customers and Google AI Ultra subscribers next.
- On Google's own chart, Argon leads on 12 of 18 tests against GPT-6 Astra and Claude Opus 5.5, VentureBeat counted.
- On the independent Artificial Analysis index, Argon scores 53, which ties GPT-6 Astra and Fable 5.1 but trails Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56.
- Its standout result is factual reliability: a 15% hallucination rate, against 51% for GPT-6 Astra.
- The introductory price is $2 and $10 per million tokens. It doubles afterward, which would erase much of Argon's cost edge.
What Is Gemini 4 Argon?
It's Google's first new top-tier model in months, and Google is pitching it hard at real work.
Vendor claim: Google's launch post calls Argon its "next era of frontier intelligence." It says the model leads on long software projects, legal and financial work, and finding and fixing security flaws.
| Detail | Gemini 4 Argon |
|---|---|
| Announced | September 30, 2026 |
| Introductory price per million tokens | $2 input, $10 output, $0.10 cached input |
| Regular price per million tokens | $4 input, $20 output |
| Maximum output per response | 1 million tokens, up from 64,000 |
| Access today | Trusted cyber defenders in the Fairwind Program |
That output limit is unusual. A million tokens in one response means Argon can, in principle, write or rewrite very large codebases in one pass. Google says it used the model internally to help migrate C and C++ code to Rust, among other projects, VentureBeat reported.
Who Can Use Gemini 4 Argon Now?
Very few people. Google is rolling it out in stages, starting with security teams.
The first users are security teams. Google calls them trusted cyber defenders, and they get Argon through its Fairwind Program. The US government gets early access too, through a voluntary pre-release process. Next come paid API customers and Google AI Ultra subscribers, then wider access "as soon as possible."
The reason is cyber capability. Google says the model is built to refuse harmful requests while still helping with legitimate security research. Starting with defenders lets them find and fix flaws before attackers get the same tool. Security firm Wiz said it used Argon to find a critical flaw in healthcare software used by hospitals, VentureBeat reported.
This mirrors a pattern across the industry. Labs now release their most capable models to security teams first. We covered the incidents behind that caution in our look at AI agent incidents at the UN Security Council.
What Does Google's Benchmark Chart Show?
A clear win on most rows, and some notable losses. Here's a selection from Google's chart, as reported by VentureBeat:
| Test | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 (software engineering) | 77.9% | 74.1% | 74.2% |
| AutomationBench | 51.3% | 41.4% | 42.5% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| Harvey Legal Agent | 19.6% | 5.4% | 3.8% |
| Terminal-Bench 4.0 | 57.4% | Not shown | 66.4% |
| FrontierSWE v2 | 55.0% | 65.5% | Not shown |
| Prompt injection success (lower is better) | 0.7% | 8.5% | 1.0% |
Vendor claim: Across 18 tests on the chart, Argon leads outright on 12 and ties on 1, VentureBeat counted. GPT-6 Astra leads on 3 and ties on 1, and Opus 5.5 leads on 2.
Notice where Argon trails. Astra wins on FrontierSWE v2, a harder software test, by more than 10 points, and Opus 5.5 wins on Terminal-Bench 4.0 by 9 points. So "best at coding" depends on which coding test you pick.
What Do Independent Tests Show for Gemini 4 Argon?
A strong model, but not the clear leader Google's chart suggests.
Independent data: On the Artificial Analysis Intelligence Index, which averages ten tests, here's where Argon lands at the High setting it tested:
| Model | Index score | Cost per task |
|---|---|---|
| Claude Opus 5.5 (max) | 58 | $5.98 |
| Claude Sonnet 5.5 (max) | 56 | $7.60 |
| Gemini 4 Argon (High, introductory price) | 53 | $1.99 |
| GPT-6 Astra (max) | 53 | $3.26 |
| Claude Fable 5.1 | 53 | Not compared here |
| GPT-6.1 Sol (max) | 52 | $0.74 |
Scores and costs come from Artificial Analysis, as reported by The Decoder and Trending Topics. Argon scored 52.6 and Astra 52.7, a tie in practice.
Against Opus 5.5, Argon trails on several individual tests in that run. It scored 57.1% on Terminal-Bench 4.0 against 59.6%, 57.1% on Humanity's Last Exam against 61.4%, and 1611 on GDPval-AA against 1846, per Trending Topics.
Where Does Gemini 4 Argon Lead?
Factual reliability, automation and human preference. These are real strengths, and some matter more in production than a composite score.
- Fewer made-up answers. On AA-Omniscience, Argon's hallucination rate is 15%. GPT-6 Astra's is 51%, and GPT-6.1 Sol's is 54%, per The Decoder. That's the lowest among models scoring above 45 on the index.
- Automation tasks. Argon leads Artificial Analysis's AutomationBench at 77.5%, about six points ahead of Claude Sonnet 5.5 at max.
- Real professional work. It tops the Vals Index at 68.9%, about two points clear of Sonnet 5.5, The Decoder reported.
- Human preference. It tops LMArena's text leaderboard at 1,525 points, where people vote blind on which of two answers they prefer.
- Prompt injection resistance. On Google's chart, attacks succeeded 0.7% of the time, the best of the three. That matters for agents that read untrusted web pages and emails.
Our view: for work where a wrong fact is costly, such as research, finance and support, a low hallucination rate can matter more than a few index points. That's Argon's strongest case.
Is Gemini 4 Argon Cheap?
For now, yes. Later, less so.
At introductory prices, Argon costs $1.99 per task on Artificial Analysis's index, about 39% less than GPT-6 Astra's $3.26 and about a third of Opus 5.5's $5.98. That's a strong price for its score.
But two things cut against that. First, Argon is wordy. It used about 62,000 output tokens per task, against about 27,000 for Astra, Trending Topics reported. Second, the price doubles to $4 and $20 after the introductory period.
Our estimate: if token use stays the same, doubling the price would put Argon near $4 per task on the same tests. That's more than Astra's $3.26 today. So budget for the regular price, not the launch price.
Caching softens the bill. At the introductory price, cached input costs $0.10 per million tokens, 95% below the normal input rate, so long prompts you send again and again cost far less.
Our guide to token efficiency explains why tokens per task drive the bill more than price per token.
How Does the Price Compare With Rivals?
At launch, Argon is priced like a mid-tier model. Later, it's priced like Opus. Here are list prices per million tokens:
| Model | Input | Output |
|---|---|---|
| Gemini 4 Argon (introductory) | $2 | $10 |
| Gemini 4 Argon (regular) | $4 | $20 |
| GPT-6.1 Sol | $2 | $10 |
| Claude Sonnet 5.5 | $2 | $10 |
| Claude Opus 5.5 | $4 | $20 |
| GPT-6 Astra | $10 | $50 |
So the introductory price matches GPT-6.1 Sol and Claude Sonnet 5.5, and the regular price matches Claude Opus 5.5. VentureBeat put it simply: the launch price is a fifth of Astra's and half of Opus 5.5's.
List price only tells part of the story, though. As the cost-per-task numbers above show, a wordy model can cost more per job than a pricier one that writes less.
Did the Gemini 4 Pro Leaks Get It Right?
Partly. In September we looked at viral claims in our piece on the Gemini 4 Pro leaks. Here's how they held up:
- Timing: mostly right. Leakers pointed to October or late September, and Google announced Argon on September 30, right at the edge of that window.
- Name: wrong. The model is called Gemini 4 Argon, not Gemini 4 Pro.
- "Beats GPT-6 Astra and Fable 5.1": half right. True on Google's chart. On Artificial Analysis's index, all three tie at 53.
- "Google achieved RSI": still unsupported. Nothing in the launch shows a self-improving system.
The lesson from that piece still applies: treat launch charts and leaks as hypotheses until you've tested them on your own work.
What Does a Million-Token Output Limit Mean?
It means one response can be enormous. Most models stop at tens of thousands of output tokens, and Argon's earlier Gemini models stopped at 64,000. A million tokens is roughly the length of several long novels.
That opens up real uses:
- Large code migrations, such as rewriting many files from one language to another in a single pass.
- Long reports and documentation that would otherwise need many stitched-together calls.
- Big data transformations, where the model writes out a full converted dataset.
It also brings costs and risks. At the introductory price, a full million-token response costs $10 in output alone. At the regular price, it's $20. Long outputs are also slow to generate and hard for a person to review.
Our view: treat the limit as a ceiling, not a target. Set a maximum output length for each task, and split huge jobs into pieces you can check.
What Should Builders Do About Gemini 4 Argon?
Prepare now, and test when you get access.
- Get in line. If you're a paid Gemini API customer, watch for access, since you're early in the queue.
- Prepare your test set. Pick 50 to 200 real tasks with known answers, drawn from the work you actually do, not from public benchmarks. Our guide to AI agent evaluation shows how.
- Test where Argon claims to shine. Fact-heavy answers, long code changes and automation tasks are its strongest cases.
- Price it at the regular rate. Plan your budget at $4 and $20 per million tokens, not the introductory $2 and $10.
- Watch the output length. A million-token output limit is powerful, but long outputs cost money. Set limits per task.
- Stay model-agnostic. Keep prompts and tools portable, so you can switch between Argon, Astra and Claude as prices and results change.
For the wider comparison, see our breakdowns of Claude Opus 5.5 and Claude Sonnet 5.5.
What Should You Watch Next?
Several open questions will decide how Argon fits into real stacks.
- Access dates. Google hasn't said when paid API customers, Ultra subscribers or everyone else get access.
- How long the launch price lasts. The regular price is double, and Google hasn't given an end date for the discount.
- Results at other settings. Artificial Analysis tested Argon at its High setting, so a higher setting, if Google offers one, could score more and cost more.
- Coding results from outside testers. Google's chart and independent runs disagree on which coding tests Argon wins.
- Rival responses. OpenAI pulled plans for GPT-6.1 Astra over safety concerns this week. Anthropic and others may answer with new models or prices.
How Van Data Team Helps Teams Test New Models Fast
We help teams evaluate new models on their own tasks within days of access. That means ready-made eval sets, cost-per-task tracking, and routing rules that send each job to the model that does it best for the money.
Gemini 4 Argon is a serious new option, especially where factual accuracy matters. If you want help testing it against your current models, our AI governance guide is a good place to start.
Article FAQ
Questions readers usually ask next.
These short answers clarify the practical follow-up questions that often come after the main article.
Need a similar system?
If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.
Book your free workflow review here.
