HomeLearn

Grok 4.7 Benchmarks: Scores, Pricing, and How It Compares

Grok 4.7 benchmarks show frontier-adjacent scores at $2/$6 pricing. See how it compares to Claude and GPT-6, and where the real cost hides.

Bhavyadeep
Written by
Bhavyadeep
Sakthy
Reviewed by
Sakthy
Last updated: 
September 22, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

Grok 4.7, released by xAI (branded SpaceXAI on its launch materials) on September 21, 2026, is a strong, cost-efficient model that sits just behind the top frontier tier. The quick read:

  • Pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6.
  • It leads on EEBench and the Harvey Legal Agent benchmark, and trails Claude Fable 5.1 on coding.
  • Its independent Artificial Analysis Intelligence Index score is 46, below Fable 5.1 and GPT-6 Astra at 53.
  • The model is slow and unusually verbose, so the real cost to finish a task runs higher than the per-token rate suggests.
  • Best fit: cost-conscious legal and engineering reasoning, rather than topping the charts on raw coding or speed.

Grok 4.7, released by xAI (branded SpaceXAI on its launch materials) on September 21, 2026, is a strong, cost-efficient model that sits just behind the top frontier tier. The quick read:

  • Pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6.
  • It leads on EEBench and the Harvey Legal Agent benchmark, and trails Claude Fable 5.1 on coding.
  • Its independent Artificial Analysis Intelligence Index score is 46, below Fable 5.1 and GPT-6 Astra at 53.
  • The model is slow and unusually verbose, so the real cost to finish a task runs higher than the per-token rate suggests.
  • Best fit: cost-conscious legal and engineering reasoning, rather than topping the charts on raw coding or speed.

Grok 4.7 posts frontier-adjacent scores at a mid-tier price

Grok 4.7 is a strong model that sits just behind the very top tier, and it costs less than most of the models it competes with. That is the short version, and the Grok 4.7 launch confirms the headline numbers.

xAI built it for coding and long-running knowledge work. The company says it works longer on hard tasks, checks its own work more carefully, and ships with the firm's best-calibrated safeguards so far. It runs at the same price and speed as the earlier Grok 4.6.

The independent picture is more measured. Artificial Analysis, which tests models on a single shared harness, places Grok 4.7 at number 16 out of 202 models on its Intelligence Index. That is well above the median of 24. It is a genuinely capable model. It is not the outright leader, and the gap between what the vendor markets and what an independent lab measures is where the interesting reading lives.

What the benchmark scores actually say

Grok 4.7 wins some categories outright and loses others by a clear margin. The pattern favors legal and engineering work over raw coding.

Two things matter before you read the table. First, these figures come from xAI's own launch post, so they carry a vendor-reported label: the maker's numbers, tested on the maker's setup, at reasoning settings that differ across models (Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, Fable 5.1 Max). Second, the DeepSWE score is a high-effort run. Throughout this article, benchmark figures are tagged one of two ways: vendor-reported (published by the model maker) or independent (measured by Artificial Analysis on a shared harness). The two are not directly comparable, so read each against its own label.

Benchmark Task type Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max
CursorBench 4.0 Software engineering 46.3% 40.4% 41.7% 51.8%
DeepSWE v1.1 Software engineering 71.0%* 65.2% 72.7% 70.0%
EEBench Electrical engineering 64.0% 53.0% 39.4% 56.4%
AA Briefcase v1.1 Multi-hour office work 1,657 1,546 1,487 1,678
Terminal-Bench 4.0 Multi-hour terminal work 38.0% 20.3% 37.3% 57.9%
Harvey Legal Agent Legal work 19.6% 15.8% 2.5% 6.7%
HealthBench Professional Clinical reasoning 56.7% 48.5% 60.5% 62.1%

High-effort score. All figures vendor-reported; settings are not identical across models, so scores are not strictly like-for-like.

Read across the rows and the story sorts itself. On EEBench, Grok 4.7 tops every model in the comparison, including Fable 5.1. On the Harvey Legal Agent benchmark it clears the rest of the field by a wide margin, with GPT-5.6 Sol scoring a striking 2.5%. Those are real strengths.

Coding is a different result. Fable 5.1 beats Grok 4.7 on CursorBench, and both Fable 5.1 and GPT-5.6 Sol edge ahead on the multi-hour Terminal-Bench task. Grok 4.7 improves sharply over Grok 4.6 on nearly every row, so the upgrade is real, but it does not take the coding crown.

The independent read adds one more data point worth holding onto. On the Artificial Analysis Intelligence Index, an independently measured composite of ten evaluations spanning agents, coding, knowledge, and scientific reasoning, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra sit higher at 53, and Claude Opus 5 lands at 51. Because it is a broad composite rather than a single coding or legal score, treat it as an overall-capability signal. Grok 4.7 shares the next tier down.

A score is only half the decision. Our best Grok alternatives guide covers what each option costs to run.

Pricing tells only half the cost story

1. The sticker price looks like a bargain

Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. That is cheaper than most frontier models on input, and far cheaper on output than Fable 5.1 at $50. On a spec sheet, it looks like a bargain.

The spec sheet leaves out how many tokens the model actually spends.

2. Verbosity pushes the real cost higher

Artificial Analysis measured Grok 4.7 generating 240 million output tokens to complete its Intelligence Index, against a median of 94 million for comparable models. The lab describes it as very verbose. Because you pay per token, that verbosity feeds straight back into the bill. Once the token spend is folded in, the independent cost-per-task figure lands at $3.74 for the xHigh tier, a benchmark-weighted estimate rather than a fixed price per user task, which is mid-pack rather than cut-rate. The same pattern shows up in broader coverage of how frontier model costs move faster than the raw pricing implies.

3. Slow output taxes long agentic tasks

Speed compounds the effect. At 39.5 output tokens per second, Grok 4.7 ranks 152nd of 202 models on the same harness. It is a slow model. For a single quick answer that barely registers. For the multi-hour agentic work xAI built it for, slow and verbose is a meaningful tax on real-world throughput.

Grok 4.7 pricing and independent performance, as of September 2026. Pricing is subject to change; verify at the source before relying on it.

Measure Grok 4.7 (xHigh) Source tier
Input price (per million tokens) $2.00 Vendor-reported
Output price (per million tokens) $6.00 Vendor-reported
Cache discount 75% Independent
Cost per Intelligence Index task (xHigh) $3.74 Independent
Output speed 39.5 tokens/sec Independent
Context window 500k tokens Independent

How Grok 4.7 compares to Claude, GPT, and Gemini

Against the current frontier, Grok 4.7 is a value pick rather than a performance leader. It wins on price-per-token and on a few specialized benchmarks. It loses on raw intelligence, speed, and coding to the models at the top.

How Grok 4.7 compares with two frontier rivals. As of September 2026. Source tier is marked per row: independent = Artificial Analysis; vendor = xAI launch post.

Measure Grok 4.7 (xHigh) Claude Fable 5.1 GPT-6 Astra Source
Intelligence Index 46 53 53 Independent
CursorBench 4.0 46.3% 51.8% Not reported Vendor-reported
Output price (per million tokens) $6.00 $50.00 Not reported Vendor-reported
Relative speed Slow Faster Faster Independent

1. Claude Fable 5.1 leads on coding and speed, at a higher price

Claude Fable 5.1 is the clearest contrast. It outscores Grok 4.7 on the Intelligence Index, on CursorBench, and on Terminal-Bench, and it is far faster on the same harness. It also costs several times more per output token. If your work is coding-heavy and budget is secondary, Fable 5.1 is the stronger tool. If cost discipline matters and your tasks lean toward legal or engineering reasoning, Grok 4.7 makes a real case.

2. GPT-6 Astra edges ahead on raw intelligence

Against OpenAI's GPT-6 Astra, Grok 4.7 again trails on the independent composite Intelligence Index, 46 to 53. The two trade wins at the benchmark level, and Grok 4.7's much lower price is its main lever.

3. Gemini pushes on speed, where Grok is weakest

For a Gemini reference point, the recent Gemini 3.8 launch shows Google pushing hard on speed, an area where Grok 4.7 is notably weak. The takeaway across all three: Grok 4.7 competes on economics, not on topping the charts.

Where Grok 4.7 still struggles

The model was built to work for hours without a human in the loop. In practice, long-horizon reliability is still its softest spot.

Independent coverage has stress-tested the "works for hours" claim and found the model often fails to hold a complex task together across a long agentic run. The published benchmarks point the same way. On the vendor-reported Terminal-Bench 4.0 comparison, Grok 4.7 scores 38.0%, a substantial limitation on that test, though the score should not be read as a claim that most real-world multi-hour tasks fail. Grok 4.6 scored 20.3% on the same test, so 4.7 is a large step forward, but it is a step from a low base.

Verbosity and speed, covered above, feed the same weakness. A model that thinks slowly and writes at length is working against itself on tasks that demand many sequential steps. The strengths are real, and so are the limits. An honest read holds both.

Grok 4.7 vs Grok 4.6: what changed

If you are already on Grok 4.6, the upgrade is worth taking. Grok 4.7 improves on nearly every published benchmark at the same $2/$6 price.

xAI trained Grok 4.7 on a new, larger base model, with a longer reinforcement learning run weighted toward tasks that take many hours. The gains are visible in the table above. Terminal-Bench nearly doubles, EEBench climbs 11 points, and the AA Briefcase and legal scores both rise. The model also natively understands the Grok Bot conversational harness, which the earlier version did not.

Since the price holds steady, there is little reason to stay on 4.6 for new work. The one caveat is the verbosity trend. If your workload is highly cost-sensitive and token-bound, test your own prompts before assuming the upgrade is free in practice.

Which builders should care

Benchmark races move fast, and today's leader rarely stays on top for a full quarter. For anyone building real software rather than testing models, what matters is having access to strong frontier models without betting the whole project on a single one.

That is the gap Emergent is built to close. Emergent turns a plain description into a working, deployable, full-stack application, so founders and operators who can describe what they need can ship it without hiring engineers. It gives you a choice of leading frontier models through a single Universal LLM Key, with unified billing, so the model landscape can shift underneath you while your build keeps moving.

The scores in this guide are a snapshot. The work of building is what lasts.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

Is Grok 4.7 better than Claude?
It depends on the task. Claude Fable 5.1 scores higher on general intelligence, coding, and speed, and it costs more. Grok 4.7 wins on price and on specialized benchmarks like EEBench and the Harvey Legal Agent test. For coding-heavy work, Claude leads. For cost-conscious legal or engineering reasoning, Grok 4.7 is competitive.
Does Grok 4.7 beat GPT-6 Astra?
Not on the headline number. GPT-6 Astra scores 53 on the Artificial Analysis Intelligence Index against Grok 4.7's 46. The two trade wins on individual benchmarks, and Grok 4.7's main advantage is its much lower per-token price rather than raw performance.
How much does Grok 4.7 cost?
Grok 4.7 is priced at $2 per million input tokens and $6 per million output tokens, as of September 2026. A faster variant runs at twice the speed for twice the price. Because the model is verbose, the real cost per completed task is higher than the per-token rate suggests, landing near $3.74 per task on independent testing.
Should I upgrade from Grok 4.6?
For most users, yes. Grok 4.7 improves on nearly every benchmark at the same price, with large gains on Terminal-Bench and EEBench. The one thing to watch is token usage. If your workload is highly cost-sensitive, test your own prompts before assuming the upgrade carries no added token cost.
Where can I access Grok 4.7?
Grok 4.7 is available through the Grok API, in Cursor, and in Grok Build, along with third-party coding harnesses and model routers. Availability is current as of September 2026 and is best confirmed at the source.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql