Grok 4.7, released by xAI (branded SpaceXAI on its launch materials) on September 21, 2026, is a strong, cost-efficient model that sits just behind the top frontier tier. The quick read:
- Pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6.
- It leads on EEBench and the Harvey Legal Agent benchmark, and trails Claude Fable 5.1 on coding.
- Its independent Artificial Analysis Intelligence Index score is 46, below Fable 5.1 and GPT-6 Astra at 53.
- The model is slow and unusually verbose, so the real cost to finish a task runs higher than the per-token rate suggests.
- Best fit: cost-conscious legal and engineering reasoning, rather than topping the charts on raw coding or speed.
Grok 4.7 posts frontier-adjacent scores at a mid-tier price
Grok 4.7 is a strong model that sits just behind the very top tier, and it costs less than most of the models it competes with. That is the short version, and the Grok 4.7 launch confirms the headline numbers.
xAI built it for coding and long-running knowledge work. The company says it works longer on hard tasks, checks its own work more carefully, and ships with the firm's best-calibrated safeguards so far. It runs at the same price and speed as the earlier Grok 4.6.
The independent picture is more measured. Artificial Analysis, which tests models on a single shared harness, places Grok 4.7 at number 16 out of 202 models on its Intelligence Index. That is well above the median of 24. It is a genuinely capable model. It is not the outright leader, and the gap between what the vendor markets and what an independent lab measures is where the interesting reading lives.
What the benchmark scores actually say
Grok 4.7 wins some categories outright and loses others by a clear margin. The pattern favors legal and engineering work over raw coding.
Two things matter before you read the table. First, these figures come from xAI's own launch post, so they carry a vendor-reported label: the maker's numbers, tested on the maker's setup, at reasoning settings that differ across models (Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, Fable 5.1 Max). Second, the DeepSWE score is a high-effort run. Throughout this article, benchmark figures are tagged one of two ways: vendor-reported (published by the model maker) or independent (measured by Artificial Analysis on a shared harness). The two are not directly comparable, so read each against its own label.
High-effort score. All figures vendor-reported; settings are not identical across models, so scores are not strictly like-for-like.
Read across the rows and the story sorts itself. On EEBench, Grok 4.7 tops every model in the comparison, including Fable 5.1. On the Harvey Legal Agent benchmark it clears the rest of the field by a wide margin, with GPT-5.6 Sol scoring a striking 2.5%. Those are real strengths.
Coding is a different result. Fable 5.1 beats Grok 4.7 on CursorBench, and both Fable 5.1 and GPT-5.6 Sol edge ahead on the multi-hour Terminal-Bench task. Grok 4.7 improves sharply over Grok 4.6 on nearly every row, so the upgrade is real, but it does not take the coding crown.
The independent read adds one more data point worth holding onto. On the Artificial Analysis Intelligence Index, an independently measured composite of ten evaluations spanning agents, coding, knowledge, and scientific reasoning, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra sit higher at 53, and Claude Opus 5 lands at 51. Because it is a broad composite rather than a single coding or legal score, treat it as an overall-capability signal. Grok 4.7 shares the next tier down.
A score is only half the decision. Our best Grok alternatives guide covers what each option costs to run.
Pricing tells only half the cost story
1. The sticker price looks like a bargain
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. That is cheaper than most frontier models on input, and far cheaper on output than Fable 5.1 at $50. On a spec sheet, it looks like a bargain.
The spec sheet leaves out how many tokens the model actually spends.
2. Verbosity pushes the real cost higher
Artificial Analysis measured Grok 4.7 generating 240 million output tokens to complete its Intelligence Index, against a median of 94 million for comparable models. The lab describes it as very verbose. Because you pay per token, that verbosity feeds straight back into the bill. Once the token spend is folded in, the independent cost-per-task figure lands at $3.74 for the xHigh tier, a benchmark-weighted estimate rather than a fixed price per user task, which is mid-pack rather than cut-rate. The same pattern shows up in broader coverage of how frontier model costs move faster than the raw pricing implies.
3. Slow output taxes long agentic tasks
Speed compounds the effect. At 39.5 output tokens per second, Grok 4.7 ranks 152nd of 202 models on the same harness. It is a slow model. For a single quick answer that barely registers. For the multi-hour agentic work xAI built it for, slow and verbose is a meaningful tax on real-world throughput.
Grok 4.7 pricing and independent performance, as of September 2026. Pricing is subject to change; verify at the source before relying on it.
How Grok 4.7 compares to Claude, GPT, and Gemini
Against the current frontier, Grok 4.7 is a value pick rather than a performance leader. It wins on price-per-token and on a few specialized benchmarks. It loses on raw intelligence, speed, and coding to the models at the top.
How Grok 4.7 compares with two frontier rivals. As of September 2026. Source tier is marked per row: independent = Artificial Analysis; vendor = xAI launch post.
1. Claude Fable 5.1 leads on coding and speed, at a higher price
Claude Fable 5.1 is the clearest contrast. It outscores Grok 4.7 on the Intelligence Index, on CursorBench, and on Terminal-Bench, and it is far faster on the same harness. It also costs several times more per output token. If your work is coding-heavy and budget is secondary, Fable 5.1 is the stronger tool. If cost discipline matters and your tasks lean toward legal or engineering reasoning, Grok 4.7 makes a real case.
2. GPT-6 Astra edges ahead on raw intelligence
Against OpenAI's GPT-6 Astra, Grok 4.7 again trails on the independent composite Intelligence Index, 46 to 53. The two trade wins at the benchmark level, and Grok 4.7's much lower price is its main lever.
3. Gemini pushes on speed, where Grok is weakest
For a Gemini reference point, the recent Gemini 3.8 launch shows Google pushing hard on speed, an area where Grok 4.7 is notably weak. The takeaway across all three: Grok 4.7 competes on economics, not on topping the charts.
Where Grok 4.7 still struggles
The model was built to work for hours without a human in the loop. In practice, long-horizon reliability is still its softest spot.
Independent coverage has stress-tested the "works for hours" claim and found the model often fails to hold a complex task together across a long agentic run. The published benchmarks point the same way. On the vendor-reported Terminal-Bench 4.0 comparison, Grok 4.7 scores 38.0%, a substantial limitation on that test, though the score should not be read as a claim that most real-world multi-hour tasks fail. Grok 4.6 scored 20.3% on the same test, so 4.7 is a large step forward, but it is a step from a low base.
Verbosity and speed, covered above, feed the same weakness. A model that thinks slowly and writes at length is working against itself on tasks that demand many sequential steps. The strengths are real, and so are the limits. An honest read holds both.
Grok 4.7 vs Grok 4.6: what changed
If you are already on Grok 4.6, the upgrade is worth taking. Grok 4.7 improves on nearly every published benchmark at the same $2/$6 price.
xAI trained Grok 4.7 on a new, larger base model, with a longer reinforcement learning run weighted toward tasks that take many hours. The gains are visible in the table above. Terminal-Bench nearly doubles, EEBench climbs 11 points, and the AA Briefcase and legal scores both rise. The model also natively understands the Grok Bot conversational harness, which the earlier version did not.
Since the price holds steady, there is little reason to stay on 4.6 for new work. The one caveat is the verbosity trend. If your workload is highly cost-sensitive and token-bound, test your own prompts before assuming the upgrade is free in practice.
Which builders should care
Benchmark races move fast, and today's leader rarely stays on top for a full quarter. For anyone building real software rather than testing models, what matters is having access to strong frontier models without betting the whole project on a single one.
That is the gap Emergent is built to close. Emergent turns a plain description into a working, deployable, full-stack application, so founders and operators who can describe what they need can ship it without hiring engineers. It gives you a choice of leading frontier models through a single Universal LLM Key, with unified billing, so the model landscape can shift underneath you while your build keeps moving.
The scores in this guide are a snapshot. The work of building is what lasts.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







