HomeLearn

Qwen 3.8 Max vs Kimi K3: Cheaper Tokens, Closer Fight

Qwen 3.8 Max vs Kimi K3: compare published benchmarks, $2/$6 vs $3/$15 pricing, and open weights. See which model wins and where the vendor claims fall short.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Published: 
Jul 22, 2026
0
 min read
Table of Contents

TL;DR

  • Qwen 3.8 Max reached general availability on August 3, 2026 with published benchmarks and a standard rate of $2 input, $6 output, and $0.25 cached input per million tokens. The preview-only, credit-only phase is over.
  • Alibaba's published table shows Qwen 3.8 Max leading PaperBench at 93.0 and posting large gains over Qwen3.7-Max, but trailing Claude Fable 5 on SWE-bench Pro, FrontierSWE, JobBench, and Humanity's Last Exam. The "second only to Fable 5" claim is not supported by Alibaba's own numbers.
  • Kimi K3 scores 57 on the Artificial Analysis Intelligence Index at launch (v4.1), currently listed at 60 on the updated v4.1.1, ranks third to fourth overall, and holds #1 on the Frontend Code Arena. Its scores are independently verified.
  • Alibaba did not benchmark Qwen 3.8 Max against Kimi K3. The two models still do not appear together on a single shared-harness table, so most side-by-side charts online quietly mix sources.
  • The one true independent head-to-head, TrilogyAI's StackPerf architecture test, scored Kimi K3 at 83 and the Qwen 3.8 preview at 80. Kimi K3 shipped open weights on July 27. Qwen's weights are promised but not yet released.


Qwen 3.8 Max vs Kimi K3 is a matchup between two frontier-scale Chinese models that landed within weeks of each other, and until recently only one of them could be judged on real numbers. That gap has closed. Alibaba took Qwen 3.8 Max to general availability on August 3, 2026 with a published benchmark table and standard per-token pricing, so both models can now be compared on specs, benchmarks, cost, and open-weight status rather than vendor claims alone. This comparison walks through each of those, flags where the published data is independently verified versus vendor-reported, and ends with a clear read on which model fits which job.

Qwen 3.8 Max and Kimi K3 at a glance

Both are multi-trillion-parameter Chinese mixture-of-experts models launched within weeks of each other, and both now have published pricing. The gaps that remain are open weights and a shared benchmark.

Spec Kimi K3 Qwen 3.8 Max
Vendor Moonshot AI Alibaba (Qwen team)
Announced July 16, 2026 July 19, 2026 (preview); August 3, 2026 (GA)
Total parameters 2.8 trillion (MoE) 2.4 trillion (MoE)
Active per token 104B (16 of 896 experts) ~95B
Context window 1,048,576 tokens 1,000,000 tokens
Max output 131K default, up to 1M 65,536 tokens
Modality Text, images, and video Text and images confirmed; broader claims unverified
Reasoning control low, high, max (default max) low, medium, xhigh (default xhigh)
Input price per 1M $3.00 $2.00
Output price per 1M $15.00 $6.00
Cached input per 1M $0.30 $0.25
Open weights Shipped July 27 under the Kimi K3 License Promised within a week of GA; not yet released
Technical report Published July 27, 2026 None published
Independent index score 57 at launch, 60 on current v4.1.1 (Artificial Analysis) Not independently indexed yet
Model ID kimi-k3 qwen3.8-max

Specs and pricing as of August 2026. Sources: Moonshot K3 blog, Qwen 3.8 announcement, Alibaba Model Studio pricing, Artificial Analysis, TrilogyAI benchmark.

Two rows carry most of the decision. Kimi K3's scores are independently verified and its weights are downloadable today. Qwen 3.8 Max is cheaper per token and posts strong vendor-reported numbers, but no third party has scored it and its weights have not shipped. Everything else is close.

What Qwen 3.8 Max's published benchmarks show

Qwen 3.8 Max posts frontier-class vendor numbers, with a standout win on research reproduction and a large jump over Qwen3.7-Max. Alibaba's announcement compares the model against Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, and Qwen3.7-Max. Every figure below is vendor-reported by Alibaba. Qwen's own scores are largely on the Claude Code harness, but the competitor scores are drawn from mixed harnesses and sources per its footnotes (Terminus 2, Codex, and official leaderboards), so the cross-model rows are not strictly apples-to-apples. Treat the numbers as vendor claims until an independent lab confirms them.

Benchmark Qwen 3.8 Max Claude Fable 5 GPT-5.6 Sol Opus 4.8 Qwen3.7-Max
PaperBench 93.0 88.8 90.5 80.3 64.8
Terminal-Bench 2.1 86.6 84.6 88.8 84.6 74.5
FrontierSWE 73.5 88.8 -- 70.0 40.7
SWE-bench Pro 67.7 80.0 64.6 69.2 60.6
JobBench 53.4 57.4 45.4 48.4 31.3
GPQA Diamond 92.6 92.6 94.1 92.0 92.4
Humanity's Last Exam 43.6 53.3 47.2 45.7 41.4

Vendor-reported by Alibaba, from its full benchmark table. Harness details in Alibaba's footnotes.

The generation jump is the clearest signal. On FrontierSWE, Qwen 3.8 Max nearly doubles its predecessor, moving from 40.7 to 73.5. On JobBench it climbs from 31.3 to 53.4. These are large gains for a single generation, and they hold across the coding and agentic suites rather than appearing on one cherry-picked test.

PaperBench is the headline win. Alibaba backs it with a demonstration where the model reproduced a research paper from nothing but the paper and a set of GPUs, wrote roughly 7,600 lines of code over about 125 hours, and then improved on the paper's own method with a 2.7-point gain on a competition math benchmark. That is the kind of long-horizon autonomy that most models cannot sustain.

The problem is the framing. Alibaba positioned Qwen 3.8 Max as "second only to Fable 5," yet its own table shows Fable 5 ahead on SWE-bench Pro, FrontierSWE, JobBench, and Humanity's Last Exam. GPT-5.6 Sol also leads on Terminal-Bench 2.1 and GPQA Diamond. Qwen 3.8 Max wins clearly on PaperBench and beats Opus 4.8 and GPT-5.6 Sol on several agentic rows. A fair reading is frontier-class and specialized, not second overall.

How Kimi K3's verified scores compare

Kimi K3 is the only model in this comparison with independently verified data, and that is its central advantage. Artificial Analysis scored it at 57 on the Intelligence Index at launch (v4.1) and currently lists it at 60 on the updated v4.1.1, ranking it third to fourth overall behind Claude Fable 5 and GPT-5.6 Sol and roughly level with Opus 4.8. No open-weight model had placed that high before. Moonshot's Kimi K3 announcement covers the launch scores and architecture in full.

Evaluation Kimi K3 Source
Intelligence Index (v4.1.1, current) 60 Artificial Analysis (verified)
Intelligence Index (v4.1, at launch) 57 Artificial Analysis (verified)
Frontend Code Arena #1 (1,679 Elo) LMArena (blind human votes)
GPQA Diamond 94% Artificial Analysis
Terminal-Bench 2.1 85% Artificial Analysis
AA-Briefcase (Elo) ~1,545 (second only to Fable 5) Artificial Analysis
Cost per task $0.94 Artificial Analysis

Scores independently verified by Artificial Analysis. The index version matters: K3 read 57 on v4.1 at launch and 60 on the current v4.1.1. Use the same version when comparing across models.

Here is the trap in every side-by-side chart you will find online, including some that look authoritative. Alibaba did not put Kimi K3 in its benchmark table, and Artificial Analysis has not yet indexed Qwen 3.8 Max. The two models have no shared-harness result. When an article shows Qwen's 86.6 on Terminal-Bench 2.1 next to K3's 85%, those come from different harnesses run by different parties, and Qwen's figure is a vendor claim while K3's is independently measured. They are not directly comparable, and anyone presenting them as a clean matchup is glossing over that.

The one real head-to-head: TrilogyAI StackPerf

The only test that ran both models on identical inputs gave Kimi K3 a three-point edge, and it used the Qwen preview endpoint rather than the GA model. TrilogyAI handed each model a frozen 269-file repository snapshot and a 60-minute window to produce an architectural analysis with evidence citations.

Metric Kimi K3 Qwen 3.8 Max Source
Final score (after factual penalties) 83/100 80/100 TrilogyAI StackPerf
Unsupported claim groups 7 7 TrilogyAI
Repository citations 274 354 TrilogyAI
Tool calls 53 44 TrilogyAI
Failed tool calls 2 (recovered) 0 TrilogyAI

Results from one matched session per model, July 19, 2026, on the Qwen preview endpoint. Source: TrilogyAI.

Kimi K3 handled revisions, regeneration, and scene history more completely, and finished faster with fewer tokens. Qwen defined cleaner system boundaries, captured stronger replay metadata, and completed every tool call without a failure. TrilogyAI's own conclusion was that the two models are complementary, and that their combined recommendation beat either one alone.

Two caveats keep this from settling the question. It is a single task type judged by a single evaluator, so it cannot establish a universal ranking. And it ran against the July preview, which Alibaba describes as a continuously updated endpoint that the GA model has since replaced. A fresh run against the shipped qwen3.8-max could move the number in either direction.

Pricing: a real comparison, finally

Qwen 3.8 Max undercuts Kimi K3 on list price across the board, though a cheaper token is not automatically a cheaper task. During the preview, Qwen was credit-only and impossible to price per token. At general availability, Alibaba published a flat standard rate.

Kimi K3 Qwen 3.8 Max
Input per 1M $3.00 $2.00
Output per 1M $15.00 $6.00
Cached input per 1M $0.30 $0.25
Context covered 1M tokens 1M tokens, flat, no long-context surcharge
Cost per task $0.94 (Artificial Analysis) Not independently measured

Kimi K3 pricing from Moonshot. Qwen 3.8 Max pricing from Alibaba Model Studio, verified against the official rate card. Pricing as on August 2026.

Qwen's output rate is the standout. At $6 per million output tokens against K3's $15, Qwen is 60 percent cheaper on the side of the bill that usually dominates for reasoning and agentic work. The flat pricing across the full 1M context also removes the long-prompt surcharge that many frontier models apply.

The catch is that list price is not the same as cost per solved task. Kimi K3 has an independently measured figure of $0.94 per task. Qwen 3.8 Max does not, because no independent lab has run it through a full cost-per-task evaluation yet. If Qwen needs more attempts or longer reasoning traces to finish a job, its cheaper tokens can still add up to a higher bill.

Want the full breakdown of what Kimi K3 costs across use cases? Read our Kimi K3 pricing guide before you commit.

Architecture and open weights

Kimi K3 has shipped its weights and a technical report; Qwen 3.8 Max has shipped neither, and that is now the sharpest difference between them. Both use mixture-of-experts designs at multi-trillion scale, but the disclosure gap is wide.

Kimi K3 activates 16 of 896 experts per token for 104 billion active parameters, using Kimi Delta Attention, Attention Residuals, and Gated MLA for key-value compression. Moonshot published full weights on Hugging Face and GitHub on July 27 under the bespoke Kimi K3 License, alongside a technical report. That license reads like a permissive open-source license through most of its length, then adds revenue-triggered conditions for large-scale commercial resellers and products above 100 million monthly active users. Internal use and access through Moonshot's own products are exempt, which covers most deployments. Read the license file before adopting.

Qwen 3.8 Max reports 2.4 trillion total parameters with roughly 95 billion active per token, which is most of what Alibaba has disclosed about its architecture. There is no technical report and no model card yet. Alibaba committed to releasing open weights within about a week of the August 3 launch, and notably promised a smaller Qwen3.8-27B alongside the flagship. A 27B-class model would run on a single rented GPU, where K3's 1.56 terabytes of weights need cluster-scale hardware. As of this writing the Qwen weights have not landed, so the model remains API-only for now.

For teams looking at the broader landscape, our Kimi K3 alternatives roundup covers six options, and the What is Kimi guide provides context on Moonshot's lineup.

When to pick Kimi K3 vs Qwen 3.8 Max

Pick Kimi K3 when you need verified data or open weights today, and Qwen 3.8 Max when list price and the newest flagship matter more than independent proof. The decision is closer than it was a month ago, because Qwen now has real pricing and real numbers behind it.

Your situation Pick Why
Need independently verified benchmarks Kimi K3 57 to 60 Intelligence Index and nine verified evaluations. Qwen's numbers are vendor-reported only.
Want the lowest token price Qwen 3.8 Max $2/$6 against K3's $3/$15, with a 60% cheaper output rate.
Need open weights you can download Kimi K3 Shipped July 27 under the Kimi K3 License. Qwen's are promised but not out.
Frontend and UI generation Kimi K3 #1 on the Frontend Code Arena. Qwen has no arena ranking.
Research reproduction and long-horizon coding Qwen 3.8 Max Leads PaperBench at 93.0 and posts the largest generation jump on agentic coding.
Architecture analysis and tool-heavy work Either (test both) StackPerf shows both are capable, K3 at 83 and Qwen at 80, with complementary strengths.
Plan to self-host later Qwen 3.8 Max The promised 27B open weight would run on a single GPU, versus K3's cluster-scale 1.56TB.
Need a forecastable cost per task Kimi K3 K3 has a measured $0.94 per task. Qwen has list price but no independent per-task figure.
Betting on "second only to Fable 5" Wait Alibaba's own table shows Fable 5 ahead on several key benchmarks.

Decision matrix based on available evidence as of August 2026. This table will change when an independent lab indexes Qwen 3.8 Max or its weights ship.

The strategic backdrop matters as much as any single row. Two Chinese labs reached the frontier within three weeks of each other, both committing to open weights at aggressive pricing. That trend holds whether or not Qwen's specific "second only to Fable 5" claim survives independent testing. The competition alone changes the economics for everyone building on frontier models.

Beyond the model comparison

If you are not building AI infrastructure and just need a working app, there is a simpler path. Emergent turns a plain-language description into a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capability without managing API keys or comparing token prices.

Skip the benchmark tables and the API setup. Describe your app and let Emergent handle the build. Start Building and see how far a prompt gets you.

Was this article helpful?
About the writer
Bhavyadeep
Bhavyadeep Sinh Rathod
Content Manager

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

Is Qwen 3.8 Max better than Kimi K3?
It depends on the workload, and neither clearly wins. Qwen 3.8 Max leads on Alibaba's vendor-reported PaperBench and costs less per token, while Kimi K3 has independently verified scores, a #1 Frontend Code Arena ranking, and a slight edge in the only head-to-head test (83 to 80). Until an independent lab scores Qwen 3.8 Max on the same harness as K3, there is no clean verdict.
How much does Qwen 3.8 Max cost compared to Kimi K3?
Qwen 3.8 Max costs $2 per million input tokens, $6 output, and $0.25 cached, on a flat rate across its full 1M-token context. Kimi K3 costs $3 input, $15 output, and $0.30 cached. Qwen is cheaper on list price across the board, with the biggest gap on output tokens. Cost per solved task can still differ, since K3 has a measured $0.94 per task and Qwen does not yet.
Did Qwen 3.8 Max's "second only to Fable 5" claim hold up?
Not on Alibaba's own benchmark table. The published numbers show Claude Fable 5 ahead of Qwen 3.8 Max on SWE-bench Pro, FrontierSWE, JobBench, and Humanity's Last Exam, and GPT-5.6 Sol ahead on Terminal-Bench 2.1 and GPQA Diamond. Qwen does lead on PaperBench and beats Opus 4.8 and GPT-5.6 Sol on several agentic rows, so frontier-class is fair, but "second overall" is not supported.
When will Qwen 3.8 Max open weights be available?
Alibaba committed to releasing open weights within about a week of the August 3 launch, including a smaller Qwen3.8-27B, but no weights had landed as of mid-August and no license has been named. Kimi K3's weights shipped July 27 under the Kimi K3 License. Both flagships require cluster-scale hardware to self-host, though the promised Qwen 27B would run on a single GPU.
Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql