This comparison is asymmetric, and that asymmetry is the most important thing to understand before reading further. Kimi K3 has independently verified benchmark scores, published per-token API pricing, and open weights you can download today. Qwen 3.8 Max has a vendor claim of being second only to Fable 5, a preview endpoint with credit-based pricing, and no published benchmarks of any kind.
Both are multi-trillion parameter Chinese MoE models launched within three days of each other in July 2026. Both target the frontier. Both committed to releasing open weights. Only K3 has shipped. The evidence available for each is fundamentally different, and this article labels that difference clearly rather than pretending both sides of the comparison sit on equal footing.
Qwen 3.8 Max and Kimi K3 at a glance
Specs as of July 2026. Sources: Moonshot K3 blog, Alibaba Qwen announcement, Coursiv integration metadata, TrilogyAI benchmark.
The cells marked "Not scored" and "Not published" reflect genuine gaps in Alibaba's public documentation. Until Alibaba publishes a model card, technical report, or benchmark table, those cells cannot be filled with anything other than speculation.
One row that used to favour Qwen no longer does. Before K3's weight release, Qwen's preview endpoint offered three reasoning settings against K3's single locked maximum. K3 now exposes low, high, and max. Neither model lets you turn thinking off entirely, so reasoning control is roughly a wash between them.
The benchmark gap: K3 has data, Qwen 3.8 does not (yet)
1. What Kimi K3's independent scores look like
K3's benchmark profile is well-established by this point. It is one of the few open-weight models with a full set of independently verified scores. Artificial Analysis scores it at 57 on the Intelligence Index, fourth overall behind Claude Fable 5 (~60), GPT-5.6 Sol (~59), and roughly on par with Claude Opus 4.8 (~56). Qwen 3.8 Max has no equivalent data. The breakdown across K3's individual evaluations:
Scores independently verified by Artificial Analysis, July 23, 2026 snapshot, as cited on Moonshot's model card. Humanity's Last Exam is reported without tool augmentation; K3 scores 56.0 with tools.
K3 also holds #1 on LMArena's Frontend Code Arena (blind developer voting) and leads vendor-reported coding benchmarks from Moonshot's launch table, including 88.3% on Terminal-Bench 2.1 (Kimi Code harness), 81.2% on FrontierSWE, and 67.5% on DeepSWE.
For a deeper look at how these scores compare to specific competitors, see our Kimi K3 vs Claude Opus 4.8 and Kimi K3 vs GPT-5.6 Sol, Terra, Luna comparisons.
2. The Qwen3.7-Max baseline (predecessor scores)
With no Qwen 3.8 benchmarks published, the predecessor Qwen3.7-Max provides the most relevant reference point. It launched in May 2026 and has been independently tested.
Qwen3.7-Max scores from Artificial Analysis and Qwen's official documentation. The Intelligence Index version matters: Qwen3.7-Max scored 56.6 on v4.0 (May 2026) but 46 on the current v4.1 (July 2026). K3's 57 is measured on v4.1. Use the same index version when comparing.
On the current Intelligence Index v4.1, Qwen3.7-Max sits at 46, which is 11 points behind K3's 57. That is a substantial gap. When the model launched in May 2026, it scored 56.6 on the then-current v4.0, placing it fifth overall. The index methodology changed between versions, which means the two numbers are not directly comparable. What matters for this comparison: on the same index version K3 was tested on, Qwen3.7-Max trails by 11 points, not 0.4.
If Qwen 3.8 truly ranks "second only to Fable 5" as Alibaba claims, it would need to jump roughly 14 points from its predecessor's current v4.1 score. That is a large leap. As AIToolsReview noted, K3 also launched with bold vendor claims that Artificial Analysis later confirmed as roughly accurate (57 vs the vendor-implied near-Fable performance). The same pattern may or may not hold for Qwen 3.8. We do not know yet.
3. The one head-to-head test (TrilogyAI StackPerf)
The only published independent comparison ran both models on the same software architecture benchmark. TrilogyAI gave both models identical frozen repository snapshots (269 files) and a 60-minute window to produce an architectural analysis with evidence citations.
Results from one matched session per model, July 19, 2026. Source: TrilogyAI.
K3 scored three points higher. It handled revisions, regeneration, and scene history more completely and finished with lower latency and fewer tokens. Qwen defined cleaner system boundaries, captured stronger replay metadata, and had zero failed tool calls.
TrilogyAI's conclusion: their agreement on the core architectural decision was stronger than either report individually, making a case for model diversity on high-value analysis work. This is one test, one task type, one evaluator. It cannot establish a universal ranking, but it does confirm that Qwen 3.8 can complete complex, tool-heavy analysis work at a level close to K3.
Pricing: published rates vs credit subscriptions
This is where a direct comparison becomes impossible. K3 has a standard, published per-token API rate. Qwen 3.8 Max does not. Any cost comparison you see elsewhere that puts a dollar-per-million-token figure next to Qwen 3.8 is either inferring from predecessor pricing or guessing from credit conversion rates that Alibaba has not published.
K3 pricing from Moonshot. Qwen pricing from Alibaba Token Plan. Qwen credit-to-token conversion is not published. Pricing as of July 2026.
K3's economics are transparent. You pay $3 per million input tokens and $15 per million output tokens, with cached input at $0.30. You can forecast your bill before running a single request.
Qwen 3.8 Max's economics are opaque. The Token Plan charges monthly subscriptions and dispenses credits, but the credit-to-token conversion rate is not publicly documented. As the eesel AI pricing analysis put it: "cheapest to try, least clear to budget." The 10% preview discount makes experimentation cheap, but production budgeting requires knowing what the standard rate will be, and Alibaba has not published one.
For teams that need to forecast AI costs across models, the Emergent Universal LLM Key handles token routing across Claude, OpenAI GPT, and Google Gemini with transparent per-model pricing.
Architecture and specs: 2.8T vs 2.4T
Both models use Mixture-of-Experts architectures at multi-trillion parameter scale. Beyond that high-level similarity, the comparison is lopsided because K3 has published architecture details and Qwen 3.8 has not. Alibaba has released no technical report, no model card, and no architecture disclosure for Qwen 3.8. Everything below about Qwen's architecture is limited to what Alibaba announced publicly or what third parties extracted from the preview endpoint.
Kimi K3 activates 16 of 896 experts per token (104 billion active), uses Kimi Delta Attention (hybrid linear attention in a 3:1 ratio with full attention), Attention Residuals, and Gated MLA for key-value compression. The published model card adds 93 layers, a 160K vocabulary, a MoonViT-V2 vision encoder at 401 million parameters, and MXFP4 weights with MXFP8 activations trained through quantization-aware methods.
Qwen 3.8 Max reports 2.4 trillion total parameters. That is the extent of confirmed architectural information. Alibaba has not disclosed the active parameter count, expert configuration, attention mechanism, or serving requirements. The model is confirmed multimodal (image input verified by TrilogyAI), and the preview endpoint supports thinking mode with low, high, and xhigh reasoning settings (xhigh is the default, per Coursiv integration metadata). Without a technical report, the architecture cannot be evaluated beyond the parameter headline.
The parameter gap (2.8T vs 2.4T) tells you the checkpoint size, not the inference cost or quality. Active parameters, caching efficiency, and provider infrastructure determine the actual economics and speed of each model.
Open weights, access, and ecosystem
K3 shipped its weights. Qwen 3.8 has not. That gap was a matter of days when this comparison first ran. It is now the single clearest difference between the two models.
Kimi K3 published full weights on Hugging Face and GitHub on July 27, 2026, alongside a technical report covering architecture, training, and evaluation. The weights ship under a bespoke document Moonshot calls the Kimi K3 License, tagged license:other on Hugging Face rather than as a standard open-source license. It reads like MIT through most of its length, then adds two commercial conditions:
- Model-as-a-Service resellers, meaning anyone giving third parties inference or fine-tuning access with control over inputs and parameters, must sign a separate agreement with Moonshot once licensee-plus-affiliate revenue passes $20 million over any consecutive 12 months
- Products above 100 million monthly active users or $20 million in monthly revenue must display "Kimi K3" in the interface
Purely internal use and access through Moonshot's own products or certified inference partners are exempt, which covers most enterprise deployments. Artificial Analysis classifies the license as requiring a separate agreement for commercial use, so it does not carry the blanket permissiveness of MIT-licensed weights from Z.ai or DeepSeek. Read the LICENSE file on the repository before making an adoption decision.
K3 remains available through the Moonshot API, Kimi.com, Kimi Work, Kimi Code, and OpenRouter, with Together AI now listed as an inference provider. Official serving recipes exist for vLLM, SGLang, and TokenSpeed.
Qwen 3.8 Max has an open-weight commitment and little else. Alibaba's X post said "going open-weight soon" without a date, a license, or a repository link. It is not clear whether the exact 2.4T Max checkpoint will be the one released, or whether a different variant will ship. The preview is available through Token Plan, Qoder (Alibaba's coding tool), and QoderWork.
The preview endpoint (qwen3.8-max-preview) receives continuous upgrades and will eventually be replaced by a formal model, per Alibaba's Token Plan documentation. That means any test run against the preview today may not reflect the model that eventually ships as the stable release.
For teams looking at the broader landscape of open-weight models, our Kimi K3 alternatives roundup covers six options, and the What is Kimi guide provides context on Moonshot's full model lineup.
When to pick Kimi K3 vs Qwen 3.8 Max
The honest answer for most decisions: pick K3 today, evaluate Qwen 3.8 when its benchmarks and pricing are published.
Decision matrix based on available evidence as of July 21, 2026. This table will change when Alibaba publishes benchmarks and pricing.
The strategic context matters too. Both models represent the same trend: Chinese labs reaching the frontier with open-weight models at aggressive pricing. As the Emerging Trajectories analysis framed it, this is "much larger than the DeepSeek moment of 2025 because it shows multiple labs can compete with and catch up to well-capitalized model vendors." That trend is real regardless of whether Qwen 3.8's specific claims hold up.
But the trend does not substitute for verified data on any individual model. Whether you pick K3 or wait for Qwen 3.8 to publish its scores, the fact that both options exist changes the economics for everyone building on frontier models.
Beyond the model comparison
If you're not building AI infrastructure and just need a working app, there's a simpler path. Emergent lets you describe an application in plain language and get a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capabilities without managing API keys or comparing token prices.
Skip the model comparisons and API setup. Describe your app and let Emergent handle the rest. Start Building and see how far a prompt gets you.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes






