No open-weight model has ever placed this high on an independent intelligence ranking. Kimi K3 landed fourth overall on the Artificial Analysis Intelligence Index within 24 hours of launch, took #1 on a blind frontend coding leaderboard, and prompted Fireworks to publish routing data showing it handles 72-96% of tasks at a fraction of the cost of the model above it. The scores are strong. The caveats are real. This article covers both.
The independent baseline: Artificial Analysis Intelligence Index
Artificial Analysis provides the cleanest cross-model comparison because it tests every model on the same harness at the same settings. Its Intelligence Index v4.1 combines nine evaluations spanning coding, agentic work, reasoning, knowledge, and scientific tasks.
All scores independently verified by Artificial Analysis, July 2026.
Artificial Analysis also runs additional evaluations outside the Intelligence Index. K3 leads two of them: AutomationBench-AA (53%, #1 among all tested models) and Harvey LAB-AA (95%, #1). It also scores well on APEX-Agents-AA (41%, second to Gemini 3.5 Flash) and EnterpriseOps-Gym-AA (45%, fourth). These are not part of the Intelligence Index composite score but provide additional signal on agentic and enterprise capabilities.
Where K3 stands out in the independent data: it ties for the top spot on τ³-Banking (33%, alongside Sol and Grok), leads on AA-LCR long-context reasoning (75%, highest tested), and scores second on SciCode (59%, behind only Fable 5). It posts strong scores on GPQA Diamond (94%, tied at the top with Sol and GPT-5.5) and GDPval-AA (59%, third behind Fable 5 and Sol).
Where K3 is weaker: CritPt physics reasoning (23%, mid-pack), Humanity's Last Exam (44%, behind Fable 5's 53% and Opus 4.8's 46%), and the hallucination rate (49% non-hallucination, significantly lower than Opus 4.8's 64% and GLM 5.2's 72%).
For head-to-head breakdowns against specific competitors, see our K3 vs Claude Opus 4.8, K3 vs GPT-5.6 Sol, Terra, Luna, and K3 vs GLM 5.2 comparisons.
AA-Briefcase: second only to Fable 5 on knowledge work
Artificial Analysis released AA-Briefcase, a new benchmark for agentic knowledge work, alongside the K3 evaluation. It tests models on realistic tasks requiring deliverables like spreadsheets, presentations, and UI mockups from complex input files.
K3 scored an Elo of 1,543, the second-highest recorded, behind only Claude Fable 5 (1,574) and ahead of GPT-5.6 Sol (1,501), Claude Sonnet 5 (1,388), and Claude Opus 4.8 (1,347). That is a +727 improvement over Kimi K2.6 (816).
Source: Artificial Analysis AA-Briefcase article, July 2026. AA reported K3's time as ~2.5x Fable's and ~3.8x Grok 4.5's (~15 min). Sol's time was not separately reported.
The quality is strong. K3's analytical quality Elo (1,754) is comparable to Fable 5 (1,744). But the operational cost is high: $10.57 per task, 83 turns per task, and nearly an hour per task on average. K3 is more verbose and slower than Fable 5 on the same work, even though the quality is close. Presentation quality (Elo 1,471) is noticeably weaker than Sol (1,660) and Opus (1,492), meaning K3's output is analytically strong but visually less polished.
Vendor-reported coding benchmarks (and the harness caveat)
Moonshot published a comparison table in its K3 technical blog. K3 leads most rows. But the models were tested on different agent harnesses, and the harness changes the score.
All scores vendor-reported by Moonshot. Competitor models in Moonshot's table were tested on their own best harnesses (Claude Code, Codex, Terminus-2).
Why the harness matters more than most readers realize
An agent benchmark does not test a model in isolation. It tests a system: the model plus its prompt, tools, retry logic, timeout, permissions, and context management. When K3 uses KimiCode and a competitor uses Claude Code or Codex, the score difference reflects the entire system, not just the model.
How large is the harness effect? NxCode documented that Claude Opus 4.8 scores 69.2% on SWE-bench through Anthropic's own scaffold and 51.9% on Scale AI's standardized SEAL board. Same model, different harness, 17.3-point gap. Gemini 3.1 Pro shows a 26.4-point spread between its verified and standardized scores.
Two datapoints from K3's own results illustrate this within a single model. On DeepSWE, K3 scores 67.5% with KimiCode and 67.3% with the common mini-SWE-agent harness. The tiny gap here is reassuring for DeepSWE specifically. But on Terminal-Bench 2.1, K3's vendor-reported 88.3% (KimiCode) compares to the Artificial Analysis independently verified 85% for K3 on their own harness. A 3.3-point gap from the same model on the same benchmark, just different harnesses.
The takeaway: treat any single-benchmark claim with context. A score labeled "Kimi K3" always means "Kimi K3 + a specific harness + specific settings." The vendor-reported table shows what Moonshot's strongest system achieves. It does not isolate the model from the infrastructure.
Frontend Code Arena: K3's standout independent result
On LMArena's Frontend Code Arena, where real developers vote blind on AI-generated website code, K3 took the #1 spot with an Elo of 1,679. It jumped 17 places from K2.6's #18 ranking and beat Claude Fable 5 in 76% of head-to-head matchups. It placed first in six of seven frontend domains.
This result is harder to game than most benchmarks. LMArena uses blind evaluation by human developers who see the output without knowing which model produced it. The sample size is large enough to be meaningful, and the result has been independently reported by multiple sources (CometAPI, AvenChat, Codersera, Techsy, wan27.org).
For teams building user-facing interfaces, landing pages, dashboards, or design-to-code workflows, this is the most practically relevant benchmark in K3's portfolio. It measures what the model actually produces when asked to build something a developer would judge, not what it scores on a standardized test.
The hallucination tradeoff
This is the benchmark story most coverage skips, and it may be the most important one for production use.
On AA-Omniscience, K3's accuracy improved from 33% (K2.6) to 46%, a 13-point gain. But the hallucination rate also climbed from 39% to 51%, a 12-point increase. The overall AA-Omniscience Index still improved (from +6 to +18) because the scoring formula rewards accuracy gains more than it penalizes hallucination increases.
What 51% hallucination rate means in practice: when K3 does not know the answer, more than half the time it guesses confidently rather than hedging or abstaining. It attempts more questions and gets more right, but it also gets more wrong.
Kili Technology's analysis explains why this happens structurally, not just at K3. A 2026 Nature paper by Kalai, Nachum, Vempala, and Zhang found that most major benchmarks use binary grading that gives zero credit for "I don't know." Under this scoring, confident guessing is the rational strategy. The authors found this pattern across HELM, the Open LLM Leaderboard, and other major evaluation suites.
This is not a K3-specific problem. Claude Fable 5 posts a 45% non-hallucination rate on AA-Omniscience (meaning it hallucinates on 55% of uncertain questions), worse than K3's 49%. The tradeoff is structural across the entire frontier. But it means benchmark accuracy improvements do not automatically translate to reliability improvements. For production use cases where a wrong answer is worse than no answer (legal, financial, medical, factual research), K3's hallucination rate is a material consideration.
For context, GLM 5.2's non-hallucination rate is 72%, and Opus 4.8's is 64%. Among frontier models, K3 is in the middle on reliability, not the worst, but not a leader.
Fireworks routing data: K3 + Fable is better than either alone
Fireworks ran one of the most practically useful third-party evaluations published so far. They tested K3 and Fable 5 across roughly 1,030 tasks in five categories (SWE, terminal, algorithmic, multi-language, and legal) using the same harness.
The finding: K3 and Fable 5 are near-tied overall but specialize differently. K3 is stronger on symbolic math, dev tooling, security, and crypto. Fable 5 wins on web/data visualization, Java, Python, and C++. On SWE tasks, K3 scored 92.4% vs Fable's 92.6%, effectively tied.
The routing insight is the actionable part. Fireworks' oracle router (which sends each task to whichever model handles it best) achieved 93% accuracy, outperforming either model alone. The router selected K3 for 72-96% of tasks because K3 was correct at lower cost on most problems, with Fable handling the long tail. K3 was up to 50x cheaper on long agentic loops and consistently lower cost across every task category.
That data supports the routing pattern recommended in our comparison articles: use K3 as the high-volume default and escalate to a premium model only on the hardest tasks. For teams exploring Kimi K3 alternatives, the Fireworks data shows the comparison is not "which model to pick" but "how to route between them."
Speed, cost, and verbosity: the operational benchmarks
The numbers that do not appear on leaderboards but determine your actual bill.
All figures from Artificial Analysis, July 2026.
K3 is notably slow (35 tok/s vs a 78 tok/s median) and notably verbose (130M output tokens across the evaluation vs a 63M median). It is also one of the most expensive models to evaluate in absolute terms ($2,710 total). But its cost per task ($0.95) is competitive because it solves hard tasks that cheaper models cannot, and it uses fewer tokens per task than its total verbosity suggests (the reasoning overhead is disproportionately concentrated on hard problems).
The practical implication: K3's economics favor hard, high-value tasks. On easy tasks, its always-on max reasoning burns tokens unnecessarily. On hard tasks, the reasoning investment pays off in solve rate. Route accordingly.
For the full pricing breakdown, see our Kimi K3 pricing deep-dive.
What the benchmarks do not tell you
Three limitations that the scores cannot capture.
Benchmarks are saturating. An ICML study of 60 text-based benchmarks found 29 at high or very high saturation (index 0.7 or above). When frontier models cluster within a few points, the difference between "winning" and "tying" may fall within measurement noise. K3's one-point edge over Opus 4.8 (57 vs 56) is real on the index but may not reflect a meaningful capability gap on your specific workload.
Lab scores do not equal deployment performance. Kili Technology's analysis documented a 37% gap between lab benchmark scores and real-world deployment performance across the industry. The tasks in your production environment are not the same tasks in the benchmark suite. Independent evaluation on your actual workload is the only test that matters for a deployment decision.
The technical report is not yet published. Moonshot has promised a technical report alongside the July 27 open-weight release. Until it arrives, training details, decontamination procedures, and ablation studies remain unknown. The benchmark scores are credible based on independent verification, but the full picture of how the model was built is still pending.
Beyond the benchmarks
If you are evaluating K3 benchmarks to decide which model to build your app on, there is a simpler path. Emergent lets you describe an application in plain language and get a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capabilities without managing API keys or comparing benchmark tables.
Skip the leaderboard research. Describe your app and let Emergent handle the rest. Start Building and see how far a prompt gets you.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes



