September 22, 2026
Grok 4.7: Same Price as Grok 4.6, but Check Cost per Task
Grok 4.7 keeps Grok 4.6's $2/$6 pricing and posts better scores. Here's what improved, what independent tests show, and why cost per task still needs a check.
Article focus
xAI launched Grok 4.7 on September 21 at the same per-token price as Grok 4.6, with big gains on its own coding and agent benchmarks. Independent scores show a smaller step up, and one early number suggests it can use far more tokens per task. Here's what to test before you switch.
Section guide
Grok 4.7 is xAI's new frontier model, released on September 21, 2026, at the same per-token price as Grok 4.6: $2 per million input tokens and $6 per million output. xAI reports big gains on coding and agent benchmarks. Independent scores show a smaller step up. And an early test suggests Grok 4.7 can use far more tokens at its highest setting, so same price per token doesn't mean same cost per task.
Key Takeaways
- xAI released Grok 4.7 on September 21, 2026, with the same per-token prices as Grok 4.6 and, xAI says, the same speed.
- On xAI's own tests, the biggest jump is Terminal-Bench 4.0, from 20.3% to 38.0%. CursorBench 4.0 rose from 40.4% to 46.3%.
- On the independent Artificial Analysis Intelligence Index, Grok 4.7 scores 46 against 44 for Grok 4.6, a real but smaller gain.
- At its highest setting, Grok 4.7 used 240 million tokens on that index, against 94 million for Grok 4.6 at high. Check cost per task, not just price per token.
- For most teams, the right move is a quick side-by-side test at the default high setting, with Grok 4.6 kept as a fallback.
What Did xAI Ship With Grok 4.7?
A stronger model at the same price. xAI's pitch is simple: better results, no price increase.
Reported fact: xAI announced Grok 4.7 on September 21, 2026. It says the model is built on "a new, larger base model" and is "served at the same price and speed as Grok 4.6." It's available in Cursor, Grok Build, the Grok API, and through third-party coding tools, model routers and cloud platforms.
According to xAI's release notes, Grok 4.7 targets coding, agent tasks and knowledge work. It takes text and image input, returns text, and has a 500,000-token context window. It offers four reasoning settings: low, medium, high and xhigh, with high as the default. xAI's models page lists its knowledge cutoff as May 2026.
xAI lists these improvements over Grok 4.6:
- A larger base model. xAI didn't publish a parameter count in its launch post.
- More reinforcement learning on a harder mix of tasks.
- Better self-checking and handling of long context.
- Better document and slide creation.
The launch also came later than first promised. Elon Musk had pointed to a mid-September release. On September 11, he said the model needed a few more days of work. According to reports, he said training had penalized response length too much, so the model gave up on hard tasks it could solve and didn't check its work enough.
We covered the previous release, including its long-context pricing and cloud VM agents, in our Grok 4.6 analysis.
What Changes Between Grok 4.6 and Grok 4.7?
Less than the version number might suggest on price and setup, and more on scores. Here's a side-by-side of what's confirmed so far, with the source for each row.
| Feature | Grok 4.6 | Grok 4.7 |
|---|---|---|
| Release date | August 12, 2026 | September 21, 2026 |
| Price per million tokens (under 200,000) | $2 input, $6 output | $2 input, $6 output |
| Context window | 500,000 tokens | 500,000 tokens |
| Artificial Analysis Intelligence Index | 44 | 46 |
| Output tokens used on that index | 94 million (high) | 240 million (xhigh) |
| Terminal-Bench 4.0 (xAI-reported) | 20.3% | 38.0% |
| Output speed | About 67 tokens per second (Artificial Analysis) | Same as Grok 4.6, per xAI; not yet measured independently |
For teams already on Grok 4.6, the upgrade is mostly a model ID change. The API, pricing and context window stay the same. What you need to check is behavior: whether the new model solves more of your tasks, and how many tokens it spends doing it.
How Much Better Is Grok 4.7 on Benchmarks?
It depends on who's measuring. xAI's own numbers show large jumps. Independent scoring shows a modest one.
Here's xAI's comparison with Grok 4.6, from its launch post. These are vendor-run results, and xAI didn't publish the full test settings. Its footnote marks the DeepSWE result for Grok 4.7 as run at high effort.
| Benchmark (xAI-reported) | Grok 4.7 | Grok 4.6 | Change |
|---|---|---|---|
| Terminal-Bench 4.0 | 38.0% | 20.3% | +17.7 points |
| EEBench | 64.0% | 53.0% | +11.0 points |
| HealthBench Professional | 56.7% | 48.5% | +8.2 points |
| CursorBench 4.0 | 46.3% | 40.4% | +5.9 points |
| DeepSWE v1.1 | 71.0% | 65.2% | +5.8 points |
| Harvey Legal Agent | 19.6% | 15.8% | +3.8 points |
The Terminal-Bench jump stands out. It measures how well a model works through tasks in a command-line environment, which is close to what coding agents do all day. Nearly doubling that score is the kind of gain that could show up in real agent work.
The independent picture is calmer. Here's how the Artificial Analysis Intelligence Index scores them:
- Grok 4.7: 46, at both high and xhigh settings.
- Grok 4.6: 44, a real but two-point gap.
- Claude Fable 5.1 and GPT-6 Astra: 53 each at their top settings, on the current leaderboard.
Grok 4.7 ranks around 15th, behind several settings of those two models and Claude Opus 5. Its pitch is value, not the top spot.
Is Grok 4.7 Really the Same Cost as Grok 4.6?
Per token, yes. Per task, maybe not. That difference matters more than the headline price.
xAI's prices are unchanged from Grok 4.6:
| Grok 4.7 pricing (per million tokens) | Prompts under 200,000 tokens | Prompts of 200,000 tokens or more |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input | $0.50 | $1.00 |
| Output | $6.00 | $12.00 |
Note the cliff at 200,000 tokens. Once a prompt crosses it, the whole request is billed at the higher rate, not just the extra part. That hasn't changed since Grok 4.6.
The bigger question is how many tokens Grok 4.7 uses. On Artificial Analysis's index, Grok 4.7 at xhigh produced 240 million output tokens. Grok 4.6 at high produced 94 million. Artificial Analysis calls Grok 4.7's figure "very verbose" next to a median of 92 million.
That's not a fair head-to-head, because the settings differ. But it's a clear warning. At $6 per million output tokens, 240 million tokens cost about $1,440, against about $564 for 94 million. Same price per token, more than twice the bill.
Here's how that plays out on a single agent task. Say a task sends 20,000 input tokens and gets 8,000 output tokens back, all under the 200,000-token line:
- At 8,000 output tokens: $0.04 for input plus $0.048 for output, about $0.09 per task.
- At 20,000 output tokens: the same $0.04 for input plus $0.12 for output, about $0.16 per task.
- At one million tasks a month: that's roughly $88,000 against $160,000, from output length alone.
These are round numbers for illustration. Your own ratio could be higher or lower, which is exactly why it's worth measuring.
There's a useful detail in the same data. Grok 4.7 scores 46 at both high and xhigh. If xhigh burns more tokens for the same score, high is the better default for most work. xAI already makes high the default.
How Fast Is the New Model?
xAI says it runs at the same speed as Grok 4.6. Independent speed data isn't in yet.
Here's what's known so far:
- xAI's claim: the same serving speed as Grok 4.6.
- Independent data: Artificial Analysis hadn't published Grok 4.7's output speed when we checked.
- The Grok 4.6 baseline: Artificial Analysis measured it at high at about 67 tokens per second, with about 43 seconds to the first token. That wait is mostly thinking time.
- The fast variant: twice the output speed at twice the price, available only through Cursor and Grok Build for now, per xAI's release notes.
Keep in mind that speed per token isn't the whole story. If Grok 4.7 writes more tokens on hard tasks, each task can take longer even at the same token rate. Measure time per finished task, not just tokens per second.
How Does It Compare With Frontier Models on Price?
It's much cheaper per token than the top models. Fable 5.1 and GPT-6 Astra both list at $10 per million input tokens and $50 per million output, as we covered in our Jev vs Fable 5.1 vs GPT-6 Astra comparison.
At list price, that means:
- Input: $2 against $10 per million tokens, five times cheaper.
- Output: $6 against $50 per million tokens, about eight times cheaper.
xAI's launch post claims it's "twice as fast, at half the price of comparable models," without naming which models it means. Treat that as marketing until you've tested it.
Cheaper per token still isn't the same as cheaper per task. Artificial Analysis shows that frontier models vary widely in how many tokens they spend.
GPT-6 Astra at high costs about $1.73 per task on its index, while Claude Fable 5.1 at high costs about $3.91. Grok 4.7's cost per task hadn't been published yet. Once it is, that's the number to compare.
For a fuller look at the two leaders, see our GPT-6 Astra vs Fable 5.1 breakdown.
What Should You Test Before You Switch?
Test on your own tasks, and measure more than accuracy. Here's a short plan:
- Pick 50 to 200 real tasks. Use work your agents already do, with known good answers.
- Run both models at high. Keep everything else the same, including prompts, tools and timeouts.
- Record four numbers per task: success, output tokens, total cost and time to finish.
- Try xhigh only where high fails. Check whether the extra tokens actually buy better answers.
- Watch the 200,000-token line. Long-context tasks that cross it cost double, so trim context where you can.
- Check behavior, not just scores. Look for giving up early, skipped checks or overly long answers.
Recall why the launch slipped: by Musk's account, the model gave up on hard tasks too early. Easing the penalty on response length may also help explain why it now writes more. Both behaviors are easy for a benchmark average to hide and easy for your own test to catch.
Our guide to AI agent evaluation covers how to build and maintain a test set like this.
Which Teams Should Use Grok 4.7?
Teams that already use Grok, and teams watching cost closely, have the most reason to test it.
- Existing Grok 4.6 users. It's a model ID change at the same price, so a trial is cheap and low-risk.
- Coding and terminal agents. The largest reported gains are in command-line and coding work.
- Cost-sensitive agent workloads. At $2 and $6 per million tokens, it's far cheaper per token than the top models.
- Cursor users. It's built into Cursor, including the fast variant.
It's a weaker fit if you need the very top of the leaderboard today. Fable 5.1 and GPT-6 Astra still score higher on independent tests. It's also a weaker fit for tasks that often exceed 200,000 tokens, where the price doubles.
Whichever you choose, wrap models behind your own interface. Switching should be a config change, because this market moves every few weeks. We made the same point about vendor risk in our piece on Grok 4.5 and the SpaceX Cursor deal.
What Do xAI's Safety Numbers Mean?
xAI reported two safety results, both from its own launch post. They're worth noting, but hard to judge without more detail.
- Biosafety. xAI says Grok 4.7 tops LatchBio's biosafety benchmark at 62.4%.
- Cyber misuse. xAI says it allows only 3.3% of risky dual-use prompts through on HackerBench v0.3.
The post doesn't explain the test methods or how other models compare on the same terms. For teams in regulated fields, ask for the full evaluation details before relying on these figures.
How Van Data Team Helps Teams Choose Between Grok 4.7 and Other Models
We help teams test new models like Grok 4.7 against their real workloads, fast. That means building labeled task sets, running models side by side, and measuring success rate, tokens, cost and time per task, not just vendor charts.
New models now land every few weeks, and each one claims to be better and cheaper. A good test harness turns each launch, including Grok 4.7, into a one-day decision instead of a guessing game. Our work on AI agent evaluation and AI agent development cost is a good place to start.
Article FAQ
Questions readers usually ask next.
These short answers clarify the practical follow-up questions that often come after the main article.
Need a similar system?
If this article maps to a workflow your team already operates, the next step is usually a scoped review of the system, constraints, and rollout path.
Book your free workflow review here.
Related articles
View allClaude Fable 5.1: Same Price, 75% Cheaper Cache

Muse Glimmer: The Reported Local Agent Model, Reviewed
Amazon Bedrock AgentCore Observability: Per-Agent Logs and Traces

