Grok 4.6 arrived on August 12, 2026, pitched as a model for long-running agents and coding. Read the Grok 4.6 benchmarks closely and a different model appears: strongest in knowledge work and research, weakest in the exact software-engineering tasks the launch centered on. It is built by xAI, now operating as SpaceXAI after SpaceX acquired the company in February 2026.
This is a full read of every published eval, what each one measures, how Grok 4.6 did, and which numbers deserve an asterisk. The 61 on the headline index tells you where the model sits on a leaderboard. It does not tell you what the model is good at, and here those two answers point in different directions.
Grok 4.6 is marketed for coding, but its benchmarks say knowledge work
Grok 4.6 takes first place on every knowledge-work benchmark in xAI's table and loses the ones that measure autonomous software engineering. That split is the clearest signal in the whole release.
It wins GDPval-AA v2 (real-world professional tasks), AA-Briefcase (long-horizon analyst work), and the Harvey legal eval. It loses DeepSWE v1.1 and Terminal-Bench v3.0, both to GPT-5.6 Sol, and neither is close. For a release whose announcement leads with codebases and agentic coding, that is a curious shape.
The read is not that Grok 4.6 is bad at code. It improved on every coding eval versus Grok 4.5. Its relative strength against frontier rivals just lives in research, analysis, and document-heavy work rather than in writing and shipping software autonomously. That distinction shapes how to read the individual numbers below.
How to read Grok 4.6's benchmarks before you trust them
Benchmark figures for Grok 4.6 come from two very different sources, and mixing them produces nonsense. Knowing which tier a number belongs to is the difference between a real comparison and a marketing screenshot.
1. xAI's launch table (self-reported)
The first tier is xAI's own launch table. It reports Grok 4.6 against Grok 4.5, GPT-5.6 Sol Max, and Fable 5 Max across 10 evals. xAI is transparent about one thing that most readers skip: the competitor figures are the best of each rival's self-reported or publicly available results. That is not a controlled four-model run on identical harnesses. It is Grok 4.6's measured scores placed next to rivals' best published numbers, which is a reasonable thing to do and a very easy thing to over-read.
2. Artificial Analysis (independently verified)
The second tier is Artificial Analysis, an independent evaluator that runs models itself on a fixed methodology. When Artificial Analysis and xAI both report a benchmark, the Artificial Analysis figure is the one to trust for cross-model comparison, because every model on its leaderboard ran through the same pipeline.
The two tiers sometimes disagree hard. The clearest case is Terminal-Bench, covered below, where the two sources report 26% and 88.4% for the same model. Both are correct. They are measuring different versions. This is exactly why a benchmark without a version number and a source is closer to decoration than data.
Source tiers for every Grok 4.6 benchmark figure
The Grok 4.6 benchmark scorecard
Here is the full launch table, with the source tier marked on each row. Bold marks the top score in each row. All competitor figures are best-of self-reported or public results, per xAI.
Grok 4.6 benchmark scorecard, source: xAI launch table (self-reported), August 2026 - competitor figures are best-of published results
Key takeaways from the scorecard:
- Grok 4.6 beats Grok 4.5 on every single row, often by wide margins, so the generational gain is real and not cherry-picked.
- Against the frontier, it wins four rows, ties one, and loses five, a strong showing but not the clean sweep a launch post implies.
The Artificial Analysis Intelligence Index: 61, and what the composite hides
Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, an independently run composite of nine benchmarks covering reasoning, coding, tool use, and knowledge. That places it sixth overall and ties it with GPT-5.6 Sol Max, one point behind Claude Fable 5 (62) and two behind Claude Opus 5 (63).
A composite score is an average, and an average smooths over the interesting parts. The 61 is the same whether a model is evenly good at everything or spiky, strong in some evals and weak in others. Grok 4.6 is spiky. Its knowledge-work components sit near the top of the field while its terminal and software-engineering components trail. The single number hides precisely the distinction that should drive your choice.
The five-point jump from Grok 4.5 (56) is worth its own note. Grok 4.5 shipped about five weeks before 4.6. Five points of index movement in five weeks, at flat pricing, is a genuine generational step rather than a version-number bump. For context on where the top of this leaderboard sits, our Claude Opus 5 review covers the model holding the number-one spot, and Opus 5 versus Fable 5 breaks down the two models directly above Grok 4.6.
Benchmark by benchmark, what each measures and how Grok 4.6 did
The scorecard is the summary. The detail is where the model's real shape shows up. Each eval below is broken into what it tests, Grok 4.6's result, and the honest read.
1. GDPval-AA v2
What it measures: agentic performance on real-world professional tasks, the multi-step knowledge work a human analyst would do, scored as an Elo rating.
Grok 4.6: 1753, the highest of the four models in xAI's table and, per Artificial Analysis, behind only Claude Opus 5 across the full field.
This is Grok 4.6's single strongest result. Artificial Analysis notes its score sits in a statistical tie with Claude Fable 5 and Qwen3.8 Max once confidence intervals overlap, so read it as frontier-tier rather than a clean solo win. Either way, real-world knowledge work is where this model is most competitive.
2. AA-Briefcase
What it measures: long-horizon agentic knowledge work, research, analysis, and document production over many steps, reported as an Elo. Artificial Analysis's private benchmark.
Grok 4.6: 1577, narrowly ahead of Fable 5 Max (1574) and comfortably ahead of GPT-5.6 Sol Max (1502).
The score matters, and the efficiency behind it matters more, which the efficiency section below covers. Artificial Analysis describes the strength as consistent across grading, presentation, and analytical quality rather than one dimension carrying the rest. For long research and analysis tasks, this is a legitimate top-tier result.
3. Harvey LAB, the outlier that needs scrutiny
What it measures: agentic legal work, scored as a criterion pass rate.
Grok 4.6: 15.8%, against 11.3% for Fable 5 Max and just 2.5% for GPT-5.6 Sol Max, a six-times gap over one of the strongest models in the table.
A gap that large is a reason to look harder, not to celebrate. When one model scores six times another on a single eval, the usual explanation is that the benchmark measures something narrow or that one model was tuned for that task shape. A 15.8% pass rate is also low in absolute terms. Grok 4.6 is better than its rivals here, and it still fails most of the legal tasks. Treat this as a directional signal, not a reason to build a legal product on one row.
4. τ³-Banking
What it measures: multi-turn customer-service interactions with tool use, the closest public benchmark to what a support agent actually does. Not in xAI's launch table.
Grok 4.6: 50.7% (per Artificial Analysis), top two alongside Qwen3.8 Max (51.3%).
Top two sounds decisive until you read the absolute number. A 50.7% score means the model handles about half of these conversations well. For a customer-facing deployment, the other half is where the risk lives, and no benchmark row fixes that. It is a strong relative result on a genuinely useful eval, with a ceiling worth remembering.
5. CursorBench v3.2
What it measures: in-editor coding assistance.
Grok 4.6: 69.9%, ahead of GPT-5.6 Sol Max (67.2%) and just behind Fable 5 Max (70.5%).
This is one of Grok 4.6's more competitive coding rows. The gap to the leader is under a point, so on everyday editor-based coding help, Grok 4.6 is squarely in the frontier pack rather than trailing it.
6. FrontierCode v1.1 (Extended)
What it measures: coding on harder, longer problems.
Grok 4.6: 61.3%, edging out GPT-5.6 Sol Max (60.6%) and sitting behind Fable 5 Max (63.6%).
Another mid-pack coding result. Grok 4.6 improved 4.7 points over Grok 4.5 here, so the coding gains are real even where the model does not lead.
7. DeepSWE v1.1, a clear loss
What it measures: autonomous software engineering, resolving real code issues end to end.
Grok 4.6: 65.9%, a wide loss to GPT-5.6 Sol Max (73%) and Fable 5 Max (70%).
This is the row that most contradicts the launch framing. Grok 4.6 improved sharply here, from Grok 4.5's 54%, so the model got much better at software engineering. It just did not catch the leaders. If autonomous coding is your primary workload, the benchmarks point you elsewhere.
Our comparison of GPT-5.6 Sol against Claude Fable 5 covers the two models that beat Grok 4.6 on this eval.
8. APEX-Agents and APEX-SWE
What they measure: long-horizon agentic tasks (APEX-Agents) and agentic software engineering specifically (APEX-SWE).
Grok 4.6: 57.5% on APEX-Agents, between GPT-5.6 Sol Max (56.7%) and Fable 5 Max (59.2%); 56.4% on APEX-SWE, behind Fable 5 Max (58.8%), with no GPT-5.6 Sol figure published.
Both are middle-of-the-pack results at the frontier. The APEX-Agents jump from Grok 4.5 (47.1%) is more than 10 points, one of the largest generational deltas in the table, and it lands exactly where xAI concentrated its agentic training.
9. Terminal-Bench, the version mismatch
What it measures: a model's ability to complete tasks in a terminal environment.
Grok 4.6: 26% on v3.0 (xAI) versus 88.4% on v2.1 (Artificial Analysis), the same model on two different versions.
This is where source tiers matter most, because the two sources disagree completely. Both numbers are real. v3.0 is the harder, newer revision, so the low score reflects a tougher test, not a broken model. The takeaway is not the specific figure. It is that a Terminal-Bench score quoted without a version tells you nothing. On the v3.0 numbers xAI published, Grok 4.6 (26%) trails GPT-5.6 Sol Max (34.6%) and Fable 5 Max (34.1%) by a wide margin, and terminal work is a real weak spot.
The efficiency benchmarks: Cost per task and turn efficiency
Grok 4.6's most underrated results are the efficiency measurements, and they come from Artificial Analysis rather than the launch table. These are benchmark outputs, not list prices, so they belong in any honest read of the scores.
Artificial Analysis measured a cost of $0.84 per task across its Intelligence Index, undercutting comparably intelligent models like GPT-5.6 Sol ($1.04). More striking is turn efficiency on long-horizon work. On AA-Briefcase, Grok 4.6 finished tasks in roughly 53 turns and about 0.5 billion input tokens on average. Claude Opus 5 took roughly 103 turns and 2 billion input tokens to reach comparable answers.
Half the turns and a quarter of the input tokens for a similar result is a real efficiency advantage on long agentic runs, where context accumulates fast and every extra turn is another billed request. That efficiency, more than any single capability score, is the practical case for the model. For how this compares against a similarly priced open-weight rival, our Kimi K3 benchmarks breakdown walks through where the two models diverge on cost and intelligence
Grok 4.6 efficiency benchmarks, source: Artificial Analysis, August 2026
The benchmarks xAI didn't put on the slide
1. The accuracy row: a 65.7% non-hallucination rate
The most important Grok 4.6 benchmark is one xAI never showed: how often the model is confidently wrong. Artificial Analysis measured it, and for anything customer-facing it outweighs every coding row.
On AA-Omniscience, Grok 4.6 scores 48.2% accuracy and a 65.7% non-hallucination rate. That second number is easy to misread. It does not mean the model hallucinates 34% of the time. It is the rate at which the model avoids inventing an answer among responses where it did not know the correct one. So when Grok 4.6 does not know something, it declines to guess about two times in three and fabricates an answer the other time.
For a coding agent, a wrong answer usually gets caught by a test. For a support agent or a research assistant, a wrong answer gets sent to a person who trusts it. A knowledge-work model with a 65.7% non-hallucination rate needs a retrieval boundary and a confidence gate around it, not raw deployment. This is not unique to Grok. Every frontier model has a version of this row. It is simply the row that should sit at the top of a benchmark writeup and almost never does.
2. Two structural caveats on the launch table
Two more caveats belong here. First, the Terminal-Bench version split already covered means any single Terminal-Bench figure is meaningless without its version. Second, xAI's competitor numbers are best-of published results, so the launch table slightly flatters Grok 4.6 against rivals whose scores were measured on different harnesses. The Artificial Analysis figures avoid that problem, which is why they anchor this article.
What practitioners measured on day one
Benchmarks describe controlled runs. Early hands-on reports add a different kind of signal, and they lined up on one theme: speed. These are individual trials, not repeatable benchmarks, so treat them as directional.
One developer ran the same new feature through both DeepSeek V4 Pro and Grok 4.6 on the same project. DeepSeek finished in about 12 minutes at 12 cents and shipped a bug. Grok 4.6 finished in a little over 3 minutes at $1.41 with a working result. That is faster and more expensive, on a single task, against a much cheaper model, which captures the actual tradeoff better than any index number.
Multiple engineers reported the same speed impression independently, several citing roughly 3x faster completion than the Claude model they had been using, with no obvious drop in engineering quality. Independent throughput data on OpenRouter backed the impression, showing real-world traffic hitting cache around 90% of the time, which pulls the effective input price well below the $2 list rate. One operational detail is worth a test before you commit: the standard API endpoint showed a 4% structured-output error rate against 0% on the zero-data-retention endpoint, which matters if you pin JSON schemas for tool calls.
What the full benchmark set says about Grok 4.6
Read end to end, the Grok 4.6 benchmarks describe a knowledge-work and research model that was announced as a coding model. It leads its launch table on GDPval-AA, AA-Briefcase, and Harvey LAB. It loses DeepSWE (65.9% to Sol's 73%) and Terminal-Bench v3.0 (26% to Sol's 34.6%). It ties GPT-5.6 Sol at 61 on the headline index while costing $0.84 per task to run, and it finishes long agentic tasks in about 53 turns against Claude Opus 5's 103.
Those numbers point to a clear verdict by workload:
- Choose Grok 4.6 for long research, analysis, and document-heavy agentic work. It leads or ties the frontier on GDPval-AA and AA-Briefcase, and its turn efficiency makes long runs cheaper than the per-token price suggests.
- Choose GPT-5.6 Sol or Fable 5 for autonomous software engineering or terminal-heavy work. Grok 4.6 improved here over Grok 4.5 but still trails both on DeepSWE and Terminal-Bench v3.0.
- Design around the 65.7% non-hallucination rate for anything customer-facing. That number, not the 61 on the leaderboard, is the one that determines whether the model is safe to put in front of users.
The one-line summary: Grok 4.6 gained five index points in five weeks at flat pricing, which is a real generational step, and it was marketed as a coding release when its measured strength is knowledge work.
Beyond the benchmarks
Benchmarks tell you how capable a model is. They do not turn that capability into software. A score of 61 says Grok 4.6 can reason and code well; it says nothing about whether the app you have in mind gets built, deployed, and run in front of real users. To cross from a capable model to a working product, you need a tool that puts these models to work, an AI app builder that takes your idea and ships it.
That is where Emergent comes in. It turns a description into a full-stack application with a real backend, real integrations, and real code you own, built by multi-agent AI rather than a single model call. Through the Universal LLM Key, Emergent runs on frontier models from Anthropic, OpenAI, and Google, so you build on the models that top these leaderboards without wiring up any of them yourself. If you have been tracking model benchmarks and want to turn that into shipped software, start building on Emergent.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes




