Kimi K3 Benchmarks: Every Score, What It Means, and Where the Caveats Are

Kimi K3 benchmark scores broken down: independent evaluations, vendor-reported coding results, the harness caveat, and what the numbers mean for real work.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Sakthyapriya Shanmugavadivel
Reviewed by
Sakthy
Published: 
Jul 23, 2026
0
 min read
Table of Contents

No open-weight model has ever placed this high on an independent intelligence ranking. Kimi K3 landed fourth overall on the Artificial Analysis Intelligence Index within 24 hours of launch, took #1 on a blind frontend coding leaderboard, and prompted Fireworks to publish routing data showing it handles 72-96% of tasks at a fraction of the cost of the model above it. The scores are strong. The caveats are real. This article covers both.

TL;DR

  • Intelligence Index: 57 (Artificial Analysis, independently verified). Fourth overall, first among open-weight models.
  • Frontend Code Arena: #1 on LMArena. Beat Claude Fable 5 in 76% of blind developer matchups. First in 6 of 7 frontend domains.
  • Coding: 88.3% on Terminal-Bench 2.1 (KimiCode harness), 81.2% on FrontierSWE, 67.5% on DeepSWE, 42.0% on SWE Marathon. All vendor-reported by Moonshot.
  • Agentic work: #2 on AA-Briefcase (Elo 1,543), behind only Fable 5. #1 on AutomationBench-AA (53%) and Harvey LAB-AA (95%).
  • Hallucination tradeoff: accuracy rose from 33% to 46% on AA-Omniscience, but hallucination rate also rose from 39% to 51%. K3 attempts more and gets more wrong.
  • The harness caveat: Moonshot's coding table mixes KimiCode, Claude Code, Codex, and mini-SWE-agent harnesses. Different harnesses can swing scores by 10-26 points.

The independent baseline: Artificial Analysis Intelligence Index

Artificial Analysis provides the cleanest cross-model comparison because it tests every model on the same harness at the same settings. Its Intelligence Index v4.1 combines nine evaluations spanning coding, agentic work, reasoning, knowledge, and scientific tasks.

Evaluation Kimi K3 Category What it tests
Intelligence Index (overall) 57 Composite Weighted score across all nine evaluations
GDPval-AA v2 59% (Elo 1,668) Agentic Economically valuable real-world tasks
τ³-Banking 33% Agentic Agentic tool use in banking workflows
Terminal-Bench v2.1 85% Coding Agentic coding and terminal operations
SciCode 59% Coding Python for scientific computing
Humanity's Last Exam 44% Reasoning Questions designed to be too hard to guess
GPQA Diamond 94% Reasoning Graduate-level scientific reasoning
CritPt 23% Reasoning Research-level physics
AA-Omniscience (accuracy) 46% Knowledge Factual accuracy across 42 topics
AA-Omniscience (non-hallucination) 49% Knowledge Rate of avoiding fabrication on uncertain questions
AA-LCR 75% Long context Long-context reasoning evaluation

All scores independently verified by Artificial Analysis, July 2026.

Artificial Analysis also runs additional evaluations outside the Intelligence Index. K3 leads two of them: AutomationBench-AA (53%, #1 among all tested models) and Harvey LAB-AA (95%, #1). It also scores well on APEX-Agents-AA (41%, second to Gemini 3.5 Flash) and EnterpriseOps-Gym-AA (45%, fourth). These are not part of the Intelligence Index composite score but provide additional signal on agentic and enterprise capabilities.

Where K3 stands out in the independent data: it ties for the top spot on τ³-Banking (33%, alongside Sol and Grok), leads on AA-LCR long-context reasoning (75%, highest tested), and scores second on SciCode (59%, behind only Fable 5). It posts strong scores on GPQA Diamond (94%, tied at the top with Sol and GPT-5.5) and GDPval-AA (59%, third behind Fable 5 and Sol).

Where K3 is weaker: CritPt physics reasoning (23%, mid-pack), Humanity's Last Exam (44%, behind Fable 5's 53% and Opus 4.8's 46%), and the hallucination rate (49% non-hallucination, significantly lower than Opus 4.8's 64% and GLM 5.2's 72%).

For head-to-head breakdowns against specific competitors, see our K3 vs Claude Opus 4.8, K3 vs GPT-5.6 Sol, Terra, Luna, and K3 vs GLM 5.2 comparisons.

AA-Briefcase: second only to Fable 5 on knowledge work

Artificial Analysis released AA-Briefcase, a new benchmark for agentic knowledge work, alongside the K3 evaluation. It tests models on realistic tasks requiring deliverables like spreadsheets, presentations, and UI mockups from complex input files.

K3 scored an Elo of 1,543, the second-highest recorded, behind only Claude Fable 5 (1,574) and ahead of GPT-5.6 Sol (1,501), Claude Sonnet 5 (1,388), and Claude Opus 4.8 (1,347). That is a +727 improvement over Kimi K2.6 (816).

Metric Kimi K3 Fable 5 Sol (max) Opus 4.8
AA-Briefcase Elo 1,543 1,574 1,501 1,347
Rubric pass rate 51% 56% 41.8% Not reported
Analytical quality Elo 1,754 1,744 Not reported Not reported
Presentation Elo 1,471 Not reported 1,660 1,492
Cost per task $10.57 Not reported Not reported Not reported
Turns per task 83 67 50 Not reported
Time per task 56.4 min ~23 min Not reported Not reported

Source: Artificial Analysis AA-Briefcase article, July 2026. AA reported K3's time as ~2.5x Fable's and ~3.8x Grok 4.5's (~15 min). Sol's time was not separately reported.

The quality is strong. K3's analytical quality Elo (1,754) is comparable to Fable 5 (1,744). But the operational cost is high: $10.57 per task, 83 turns per task, and nearly an hour per task on average. K3 is more verbose and slower than Fable 5 on the same work, even though the quality is close. Presentation quality (Elo 1,471) is noticeably weaker than Sol (1,660) and Opus (1,492), meaning K3's output is analytically strong but visually less polished.

Vendor-reported coding benchmarks (and the harness caveat)

Moonshot published a comparison table in its K3 technical blog. K3 leads most rows. But the models were tested on different agent harnesses, and the harness changes the score.

Benchmark K3 K3 harness What it tests
Terminal-Bench 2.1 88.3% KimiCode Verified terminal tasks in isolated environments
FrontierSWE 81.2% KimiCode Extremely difficult implementation and research tasks
DeepSWE 67.5% KimiCode (67.3% on mini-SWE-agent) Fresh, contamination-resistant repository issues
SWE Marathon 42.0% Claude Code Multi-hour, whole-project engineering work (20 tasks)
Program Bench 77.8% Internal Rebuild CLI programs from binary behavior (raw pass rate, not fully resolved)
Automation Bench 30.8% Internal Autonomous SaaS workflow tasks
BrowseComp 91.2% Context compaction at 300K Deep web research (90.4% without compaction)

All scores vendor-reported by Moonshot. Competitor models in Moonshot's table were tested on their own best harnesses (Claude Code, Codex, Terminus-2).

Why the harness matters more than most readers realize

An agent benchmark does not test a model in isolation. It tests a system: the model plus its prompt, tools, retry logic, timeout, permissions, and context management. When K3 uses KimiCode and a competitor uses Claude Code or Codex, the score difference reflects the entire system, not just the model.

How large is the harness effect? NxCode documented that Claude Opus 4.8 scores 69.2% on SWE-bench through Anthropic's own scaffold and 51.9% on Scale AI's standardized SEAL board. Same model, different harness, 17.3-point gap. Gemini 3.1 Pro shows a 26.4-point spread between its verified and standardized scores.

Two datapoints from K3's own results illustrate this within a single model. On DeepSWE, K3 scores 67.5% with KimiCode and 67.3% with the common mini-SWE-agent harness. The tiny gap here is reassuring for DeepSWE specifically. But on Terminal-Bench 2.1, K3's vendor-reported 88.3% (KimiCode) compares to the Artificial Analysis independently verified 85% for K3 on their own harness. A 3.3-point gap from the same model on the same benchmark, just different harnesses.

The takeaway: treat any single-benchmark claim with context. A score labeled "Kimi K3" always means "Kimi K3 + a specific harness + specific settings." The vendor-reported table shows what Moonshot's strongest system achieves. It does not isolate the model from the infrastructure.

Frontend Code Arena: K3's standout independent result

On LMArena's Frontend Code Arena, where real developers vote blind on AI-generated website code, K3 took the #1 spot with an Elo of 1,679. It jumped 17 places from K2.6's #18 ranking and beat Claude Fable 5 in 76% of head-to-head matchups. It placed first in six of seven frontend domains.

This result is harder to game than most benchmarks. LMArena uses blind evaluation by human developers who see the output without knowing which model produced it. The sample size is large enough to be meaningful, and the result has been independently reported by multiple sources (CometAPI, AvenChat, Codersera, Techsy, wan27.org).

For teams building user-facing interfaces, landing pages, dashboards, or design-to-code workflows, this is the most practically relevant benchmark in K3's portfolio. It measures what the model actually produces when asked to build something a developer would judge, not what it scores on a standardized test.

The hallucination tradeoff

This is the benchmark story most coverage skips, and it may be the most important one for production use.

On AA-Omniscience, K3's accuracy improved from 33% (K2.6) to 46%, a 13-point gain. But the hallucination rate also climbed from 39% to 51%, a 12-point increase. The overall AA-Omniscience Index still improved (from +6 to +18) because the scoring formula rewards accuracy gains more than it penalizes hallucination increases.

What 51% hallucination rate means in practice: when K3 does not know the answer, more than half the time it guesses confidently rather than hedging or abstaining. It attempts more questions and gets more right, but it also gets more wrong.

Kili Technology's analysis explains why this happens structurally, not just at K3. A 2026 Nature paper by Kalai, Nachum, Vempala, and Zhang found that most major benchmarks use binary grading that gives zero credit for "I don't know." Under this scoring, confident guessing is the rational strategy. The authors found this pattern across HELM, the Open LLM Leaderboard, and other major evaluation suites.

This is not a K3-specific problem. Claude Fable 5 posts a 45% non-hallucination rate on AA-Omniscience (meaning it hallucinates on 55% of uncertain questions), worse than K3's 49%. The tradeoff is structural across the entire frontier. But it means benchmark accuracy improvements do not automatically translate to reliability improvements. For production use cases where a wrong answer is worse than no answer (legal, financial, medical, factual research), K3's hallucination rate is a material consideration.

For context, GLM 5.2's non-hallucination rate is 72%, and Opus 4.8's is 64%. Among frontier models, K3 is in the middle on reliability, not the worst, but not a leader.

Fireworks routing data: K3 + Fable is better than either alone

Fireworks ran one of the most practically useful third-party evaluations published so far. They tested K3 and Fable 5 across roughly 1,030 tasks in five categories (SWE, terminal, algorithmic, multi-language, and legal) using the same harness.

The finding: K3 and Fable 5 are near-tied overall but specialize differently. K3 is stronger on symbolic math, dev tooling, security, and crypto. Fable 5 wins on web/data visualization, Java, Python, and C++. On SWE tasks, K3 scored 92.4% vs Fable's 92.6%, effectively tied.

The routing insight is the actionable part. Fireworks' oracle router (which sends each task to whichever model handles it best) achieved 93% accuracy, outperforming either model alone. The router selected K3 for 72-96% of tasks because K3 was correct at lower cost on most problems, with Fable handling the long tail. K3 was up to 50x cheaper on long agentic loops and consistently lower cost across every task category.

That data supports the routing pattern recommended in our comparison articles: use K3 as the high-volume default and escalate to a premium model only on the hardest tasks. For teams exploring Kimi K3 alternatives, the Fireworks data shows the comparison is not "which model to pick" but "how to route between them."

Speed, cost, and verbosity: the operational benchmarks

The numbers that do not appear on leaderboards but determine your actual bill.

Metric Kimi K3 Field context Source
Output speed 35 tok/s Median ~78 tok/s for comparable models Artificial Analysis
Time to first token 4.78s Median ~2.77s Artificial Analysis
Cost per Intelligence Index task $0.95 Cheaper than Sol ($1.04), Fable ($2.75). More than GLM ($0.47) Artificial Analysis
Total evaluation cost $2,710 Artificial Analysis
Output tokens per task ~24K (18K reasoning + 6K answer) Sol: ~15K, Opus: ~41K, GLM: ~43K Artificial Analysis
Total output tokens across full eval 130M Median: 63M among comparable models Artificial Analysis
AA-Briefcase time per task 56.4 min Fable: ~23 min, Grok: ~15 min Artificial Analysis

All figures from Artificial Analysis, July 2026.

K3 is notably slow (35 tok/s vs a 78 tok/s median) and notably verbose (130M output tokens across the evaluation vs a 63M median). It is also one of the most expensive models to evaluate in absolute terms ($2,710 total). But its cost per task ($0.95) is competitive because it solves hard tasks that cheaper models cannot, and it uses fewer tokens per task than its total verbosity suggests (the reasoning overhead is disproportionately concentrated on hard problems).

The practical implication: K3's economics favor hard, high-value tasks. On easy tasks, its always-on max reasoning burns tokens unnecessarily. On hard tasks, the reasoning investment pays off in solve rate. Route accordingly.

For the full pricing breakdown, see our Kimi K3 pricing deep-dive.

What the benchmarks do not tell you

Three limitations that the scores cannot capture.

Benchmarks are saturating. An ICML study of 60 text-based benchmarks found 29 at high or very high saturation (index 0.7 or above). When frontier models cluster within a few points, the difference between "winning" and "tying" may fall within measurement noise. K3's one-point edge over Opus 4.8 (57 vs 56) is real on the index but may not reflect a meaningful capability gap on your specific workload.

Lab scores do not equal deployment performance. Kili Technology's analysis documented a 37% gap between lab benchmark scores and real-world deployment performance across the industry. The tasks in your production environment are not the same tasks in the benchmark suite. Independent evaluation on your actual workload is the only test that matters for a deployment decision.

The technical report is not yet published. Moonshot has promised a technical report alongside the July 27 open-weight release. Until it arrives, training details, decontamination procedures, and ablation studies remain unknown. The benchmark scores are credible based on independent verification, but the full picture of how the model was built is still pending.

Beyond the benchmarks

If you are evaluating K3 benchmarks to decide which model to build your app on, there is a simpler path. Emergent lets you describe an application in plain language and get a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capabilities without managing API keys or comparing benchmark tables.

Skip the leaderboard research. Describe your app and let Emergent handle the rest. Start Building and see how far a prompt gets you.

Was this article helpful?
About the writer
Bhavyadeep
Bhavyadeep Sinh Rathod
Content Manager

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

What is Kimi K3's Intelligence Index score?
57 on Artificial Analysis Intelligence Index v4.1, placing it fourth overall behind Claude Fable 5 (~60), GPT-5.6 Sol (~59), and roughly on par with Claude Opus 4.8 (~56). It is the highest-scoring open-weight model on this index.
Is Kimi K3 better than Claude Fable 5?
On most independent benchmarks, no. Fable 5 leads the Intelligence Index (60 vs 57) and wins on FrontierSWE and Humanity's Last Exam. K3 beats Fable 5 on Frontend Code Arena (#1 vs #2), AA-Briefcase analytical quality (comparable Elos), and SWE Marathon (42.0 vs 35.0). Fireworks' routing data shows them near-tied overall at 92.4% vs 92.6% on SWE tasks.
How reliable are K3's benchmark scores?
The Artificial Analysis scores are independently verified and reliable. The vendor-reported coding scores from Moonshot's table use different agent harnesses per model, which can swing scores by 10-26 points. Treat vendor-reported scores as directional evidence of system-level performance, not definitive model-level rankings. Independent reproduction on common harnesses will strengthen the picture.
Does Kimi K3 hallucinate a lot?
K3's non-hallucination rate on AA-Omniscience is 49%, meaning it fabricates answers on 51% of uncertain questions. That is worse than GLM 5.2 (72% non-hallucination) and Opus 4.8 (64%) but better than Fable 5 (45%). The tradeoff is structural across all frontier models: benchmark scoring rewards confident answers and penalizes abstention.
What is the Frontend Code Arena and why does K3 lead it?
LMArena's Frontend Code Arena is a blind evaluation where real developers vote on AI-generated website code without knowing which model produced it. K3 ranked #1 with 1,679 Elo, beating Fable 5 in 76% of matchups and placing first in 6 of 7 frontend domains. It is one of the strongest independent signals for practical frontend coding quality.
Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql