Qwen 3.8 Max vs Kimi K3 is a matchup between two frontier-scale Chinese models that landed within weeks of each other, and until recently only one of them could be judged on real numbers. That gap has closed. Alibaba took Qwen 3.8 Max to general availability on August 3, 2026 with a published benchmark table and standard per-token pricing, so both models can now be compared on specs, benchmarks, cost, and open-weight status rather than vendor claims alone. This comparison walks through each of those, flags where the published data is independently verified versus vendor-reported, and ends with a clear read on which model fits which job.
Qwen 3.8 Max and Kimi K3 at a glance
Both are multi-trillion-parameter Chinese mixture-of-experts models launched within weeks of each other, and both now have published pricing. The gaps that remain are open weights and a shared benchmark.
Specs and pricing as of August 2026. Sources: Moonshot K3 blog, Qwen 3.8 announcement, Alibaba Model Studio pricing, Artificial Analysis, TrilogyAI benchmark.
Two rows carry most of the decision. Kimi K3's scores are independently verified and its weights are downloadable today. Qwen 3.8 Max is cheaper per token and posts strong vendor-reported numbers, but no third party has scored it and its weights have not shipped. Everything else is close.
What Qwen 3.8 Max's published benchmarks show
Qwen 3.8 Max posts frontier-class vendor numbers, with a standout win on research reproduction and a large jump over Qwen3.7-Max. Alibaba's announcement compares the model against Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, and Qwen3.7-Max. Every figure below is vendor-reported by Alibaba. Qwen's own scores are largely on the Claude Code harness, but the competitor scores are drawn from mixed harnesses and sources per its footnotes (Terminus 2, Codex, and official leaderboards), so the cross-model rows are not strictly apples-to-apples. Treat the numbers as vendor claims until an independent lab confirms them.
Vendor-reported by Alibaba, from its full benchmark table. Harness details in Alibaba's footnotes.
The generation jump is the clearest signal. On FrontierSWE, Qwen 3.8 Max nearly doubles its predecessor, moving from 40.7 to 73.5. On JobBench it climbs from 31.3 to 53.4. These are large gains for a single generation, and they hold across the coding and agentic suites rather than appearing on one cherry-picked test.
PaperBench is the headline win. Alibaba backs it with a demonstration where the model reproduced a research paper from nothing but the paper and a set of GPUs, wrote roughly 7,600 lines of code over about 125 hours, and then improved on the paper's own method with a 2.7-point gain on a competition math benchmark. That is the kind of long-horizon autonomy that most models cannot sustain.
The problem is the framing. Alibaba positioned Qwen 3.8 Max as "second only to Fable 5," yet its own table shows Fable 5 ahead on SWE-bench Pro, FrontierSWE, JobBench, and Humanity's Last Exam. GPT-5.6 Sol also leads on Terminal-Bench 2.1 and GPQA Diamond. Qwen 3.8 Max wins clearly on PaperBench and beats Opus 4.8 and GPT-5.6 Sol on several agentic rows. A fair reading is frontier-class and specialized, not second overall.
How Kimi K3's verified scores compare
Kimi K3 is the only model in this comparison with independently verified data, and that is its central advantage. Artificial Analysis scored it at 57 on the Intelligence Index at launch (v4.1) and currently lists it at 60 on the updated v4.1.1, ranking it third to fourth overall behind Claude Fable 5 and GPT-5.6 Sol and roughly level with Opus 4.8. No open-weight model had placed that high before. Moonshot's Kimi K3 announcement covers the launch scores and architecture in full.
Scores independently verified by Artificial Analysis. The index version matters: K3 read 57 on v4.1 at launch and 60 on the current v4.1.1. Use the same version when comparing across models.
Here is the trap in every side-by-side chart you will find online, including some that look authoritative. Alibaba did not put Kimi K3 in its benchmark table, and Artificial Analysis has not yet indexed Qwen 3.8 Max. The two models have no shared-harness result. When an article shows Qwen's 86.6 on Terminal-Bench 2.1 next to K3's 85%, those come from different harnesses run by different parties, and Qwen's figure is a vendor claim while K3's is independently measured. They are not directly comparable, and anyone presenting them as a clean matchup is glossing over that.
The one real head-to-head: TrilogyAI StackPerf
The only test that ran both models on identical inputs gave Kimi K3 a three-point edge, and it used the Qwen preview endpoint rather than the GA model. TrilogyAI handed each model a frozen 269-file repository snapshot and a 60-minute window to produce an architectural analysis with evidence citations.
Results from one matched session per model, July 19, 2026, on the Qwen preview endpoint. Source: TrilogyAI.
Kimi K3 handled revisions, regeneration, and scene history more completely, and finished faster with fewer tokens. Qwen defined cleaner system boundaries, captured stronger replay metadata, and completed every tool call without a failure. TrilogyAI's own conclusion was that the two models are complementary, and that their combined recommendation beat either one alone.
Two caveats keep this from settling the question. It is a single task type judged by a single evaluator, so it cannot establish a universal ranking. And it ran against the July preview, which Alibaba describes as a continuously updated endpoint that the GA model has since replaced. A fresh run against the shipped qwen3.8-max could move the number in either direction.
Pricing: a real comparison, finally
Qwen 3.8 Max undercuts Kimi K3 on list price across the board, though a cheaper token is not automatically a cheaper task. During the preview, Qwen was credit-only and impossible to price per token. At general availability, Alibaba published a flat standard rate.
Kimi K3 pricing from Moonshot. Qwen 3.8 Max pricing from Alibaba Model Studio, verified against the official rate card. Pricing as on August 2026.
Qwen's output rate is the standout. At $6 per million output tokens against K3's $15, Qwen is 60 percent cheaper on the side of the bill that usually dominates for reasoning and agentic work. The flat pricing across the full 1M context also removes the long-prompt surcharge that many frontier models apply.
The catch is that list price is not the same as cost per solved task. Kimi K3 has an independently measured figure of $0.94 per task. Qwen 3.8 Max does not, because no independent lab has run it through a full cost-per-task evaluation yet. If Qwen needs more attempts or longer reasoning traces to finish a job, its cheaper tokens can still add up to a higher bill.
Want the full breakdown of what Kimi K3 costs across use cases? Read our Kimi K3 pricing guide before you commit.
Architecture and open weights
Kimi K3 has shipped its weights and a technical report; Qwen 3.8 Max has shipped neither, and that is now the sharpest difference between them. Both use mixture-of-experts designs at multi-trillion scale, but the disclosure gap is wide.
Kimi K3 activates 16 of 896 experts per token for 104 billion active parameters, using Kimi Delta Attention, Attention Residuals, and Gated MLA for key-value compression. Moonshot published full weights on Hugging Face and GitHub on July 27 under the bespoke Kimi K3 License, alongside a technical report. That license reads like a permissive open-source license through most of its length, then adds revenue-triggered conditions for large-scale commercial resellers and products above 100 million monthly active users. Internal use and access through Moonshot's own products are exempt, which covers most deployments. Read the license file before adopting.
Qwen 3.8 Max reports 2.4 trillion total parameters with roughly 95 billion active per token, which is most of what Alibaba has disclosed about its architecture. There is no technical report and no model card yet. Alibaba committed to releasing open weights within about a week of the August 3 launch, and notably promised a smaller Qwen3.8-27B alongside the flagship. A 27B-class model would run on a single rented GPU, where K3's 1.56 terabytes of weights need cluster-scale hardware. As of this writing the Qwen weights have not landed, so the model remains API-only for now.
For teams looking at the broader landscape, our Kimi K3 alternatives roundup covers six options, and the What is Kimi guide provides context on Moonshot's lineup.
When to pick Kimi K3 vs Qwen 3.8 Max
Pick Kimi K3 when you need verified data or open weights today, and Qwen 3.8 Max when list price and the newest flagship matter more than independent proof. The decision is closer than it was a month ago, because Qwen now has real pricing and real numbers behind it.
Decision matrix based on available evidence as of August 2026. This table will change when an independent lab indexes Qwen 3.8 Max or its weights ship.
The strategic backdrop matters as much as any single row. Two Chinese labs reached the frontier within three weeks of each other, both committing to open weights at aggressive pricing. That trend holds whether or not Qwen's specific "second only to Fable 5" claim survives independent testing. The competition alone changes the economics for everyone building on frontier models.
Beyond the model comparison
If you are not building AI infrastructure and just need a working app, there is a simpler path. Emergent turns a plain-language description into a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capability without managing API keys or comparing token prices.
Skip the benchmark tables and the API setup. Describe your app and let Emergent handle the build. Start Building and see how far a prompt gets you.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







