This comparison is asymmetric, and that asymmetry is the most important thing to understand before reading further. Kimi K3 has independently verified benchmark scores, published per-token API pricing, and a confirmed open-weight release date. Qwen 3.8 Max has a vendor claim of being second only to Fable 5, a preview endpoint with credit-based pricing, and no published benchmarks of any kind.
Both are multi-trillion parameter Chinese MoE models launched within three days of each other in July 2026. Both target the frontier. Both have committed to releasing open weights, though only K3 has a confirmed date. The evidence available for each is fundamentally different, and this article labels that difference clearly rather than pretending both sides of the comparison sit on equal footing.
Qwen 3.8 Max and Kimi K3 at a glance
Specs as of July 2026. Sources: Moonshot K3 blog, Alibaba Qwen announcement, Coursiv integration metadata, TrilogyAI benchmark.
The cells marked "Not scored" and "Not published" reflect genuine gaps in Alibaba's public documentation. Until Alibaba publishes a model card, technical report, or benchmark table, those cells cannot be filled with anything other than speculation.
The benchmark gap: K3 has data, Qwen 3.8 does not (yet)
1. What Kimi K3's independent scores look like
K3's benchmark profile is well-established by this point. It is one of the few open-weight models with a full set of independently verified scores. Artificial Analysis scores it at 57 on the Intelligence Index, fourth overall behind Claude Fable 5 (~60), GPT-5.6 Sol (~59), and roughly on par with Claude Opus 4.8 (~56). Qwen 3.8 Max has no equivalent data. The breakdown across K3's individual evaluations:
All scores independently verified by Artificial Analysis, July 2026.
K3 also holds #1 on LMArena's Frontend Code Arena (blind developer voting) and leads vendor-reported coding benchmarks from Moonshot's launch table, including 88.3% on Terminal-Bench 2.1 (KimiCode harness), 81.2% on FrontierSWE, and 67.5% on DeepSWE.
For a deeper look at how these scores compare to specific competitors, see our Kimi K3 vs Claude Opus 4.8 and Kimi K3 vs GPT-5.6 Sol, Terra, Luna comparisons.
2. The Qwen3.7-Max baseline (predecessor scores)
With no Qwen 3.8 benchmarks published, the predecessor Qwen3.7-Max provides the most relevant reference point. It launched in May 2026 and has been independently tested.
Qwen3.7-Max scores from Artificial Analysis and Qwen's official documentation. The Intelligence Index version matters: Qwen3.7-Max scored 56.6 on v4.0 (May 2026) but 46 on the current v4.1 (July 2026). K3's 57 is measured on v4.1. Use the same index version when comparing.
On the current Intelligence Index v4.1, Qwen3.7-Max sits at 46, which is 11 points behind K3's 57. That is a substantial gap. When the model launched in May 2026, it scored 56.6 on the then-current v4.0, placing it fifth overall. The index methodology changed between versions, which means the two numbers are not directly comparable. What matters for this comparison: on the same index version K3 was tested on, Qwen3.7-Max trails by 11 points, not 0.4.
If Qwen 3.8 truly ranks "second only to Fable 5" as Alibaba claims, it would need to jump roughly 14 points from its predecessor's current v4.1 score. That is a large leap. As AIToolsReview noted, K3 also launched with bold vendor claims that Artificial Analysis later confirmed as roughly accurate (57 vs the vendor-implied near-Fable performance). The same pattern may or may not hold for Qwen 3.8. We do not know yet.
3. The one head-to-head test (TrilogyAI StackPerf)
The only published independent comparison ran both models on the same software architecture benchmark. TrilogyAI gave both models identical frozen repository snapshots (269 files) and a 60-minute window to produce an architectural analysis with evidence citations.
Results from one matched session per model, July 19, 2026. Source: TrilogyAI.
K3 scored three points higher. It handled revisions, regeneration, and scene history more completely and finished with lower latency and fewer tokens. Qwen defined cleaner system boundaries, captured stronger replay metadata, and had zero failed tool calls.
TrilogyAI's conclusion: their agreement on the core architectural decision was stronger than either report individually, making a case for model diversity on high-value analysis work. This is one test, one task type, one evaluator. It cannot establish a universal ranking, but it does confirm that Qwen 3.8 can complete complex, tool-heavy analysis work at a level close to K3.
Pricing: published rates vs credit subscriptions
This is where a direct comparison becomes impossible. K3 has a standard, published per-token API rate. Qwen 3.8 Max does not. Any cost comparison you see elsewhere that puts a dollar-per-million-token figure next to Qwen 3.8 is either inferring from predecessor pricing or guessing from credit conversion rates that Alibaba has not published.
K3 pricing from Moonshot. Qwen pricing from Alibaba Token Plan. Qwen credit-to-token conversion is not published.
K3's economics are transparent. You pay $3 per million input tokens and $15 per million output tokens, with cached input at $0.30. You can forecast your bill before running a single request.
Qwen 3.8 Max's economics are opaque. The Token Plan charges monthly subscriptions and dispenses credits, but the credit-to-token conversion rate is not publicly documented. As the eesel AI pricing analysis put it: "cheapest to try, least clear to budget." The 10% preview discount makes experimentation cheap, but production budgeting requires knowing what the standard rate will be, and Alibaba has not published one.
For teams that need to forecast AI costs across models, the Emergent Universal LLM Key handles token routing across Claude, OpenAI GPT, and Google Gemini with transparent per-model pricing.
Architecture and specs: 2.8T vs 2.4T
Both models use Mixture-of-Experts architectures at multi-trillion parameter scale. Beyond that high-level similarity, the comparison is lopsided because K3 has published architecture details and Qwen 3.8 has not. Alibaba has released no technical report, no model card, and no architecture disclosure for Qwen 3.8. Everything below about Qwen's architecture is limited to what Alibaba announced publicly or what third parties extracted from the preview endpoint.
Kimi K3 activates 16 of 896 experts per token (~50B active), uses Kimi Delta Attention (hybrid linear attention in a 3:1 ratio with full attention), Attention Residuals, and Gated MLA for key-value compression. Moonshot published these details in its technical blog and recommends 64+ accelerators for self-hosting. The model footprint is ~1.4 TB before KV cache.
Qwen 3.8 Max reports 2.4 trillion total parameters. That is the extent of confirmed architectural information. Alibaba has not disclosed the active parameter count, expert configuration, attention mechanism, or serving requirements. The model is confirmed multimodal (image input verified by TrilogyAI), and the preview endpoint supports thinking mode with low, high, and xhigh reasoning settings (xhigh is the default, per Coursiv integration metadata). Without a technical report, the architecture cannot be evaluated beyond the parameter headline.
The parameter gap (2.8T vs 2.4T) tells you the checkpoint size, not the inference cost or quality. Active parameters, caching efficiency, and provider infrastructure determine the actual economics and speed of each model.
Open weights, access, and ecosystem
Both models promise open weights. Neither is fully open today. But the gap between "promise" and "confirmed" is wide, and it matters if you are making infrastructure decisions.
Kimi K3 has a confirmed open-weight release on July 27, 2026 under a Modified MIT license (the single attribution clause triggers only above 100 million MAU). Moonshot published the date, the license terms, and the deployment recommendations (64+ accelerators). Until then, K3 is available through the Moonshot API, Kimi.com, Kimi Work, Kimi Code, and OpenRouter. Standard API pricing applies across all channels.
Qwen 3.8 Max has an open-weight commitment with no date, no license, no file location, and no confirmed checkpoint details. Alibaba's X post said "going open-weight soon" without further specifics. It is not clear whether the exact 2.4T Max checkpoint will be the one released, or whether a different variant will ship. The preview is available through Token Plan, Qoder (Alibaba's coding tool), and QoderWork.
The preview endpoint (qwen3.8-max-preview) receives continuous upgrades and will eventually be replaced by a formal model, per Alibaba's Token Plan documentation. That means any test run against the preview today may not reflect the model that eventually ships as the stable release.
For teams looking at the broader landscape of open-weight models, our Kimi K3 alternatives roundup covers six options, and the What is Kimi guide provides context on Moonshot's full model lineup.
When to pick Kimi K3 vs Qwen 3.8 Max
The honest answer for most decisions: pick K3 today, evaluate Qwen 3.8 when its benchmarks and pricing are published.
Decision matrix based on available evidence as of July 21, 2026. This table will change when Alibaba publishes benchmarks and pricing.
The strategic context matters too. Both models represent the same trend: Chinese labs reaching the frontier with open-weight models at aggressive pricing. As the Emerging Trajectories analysis framed it, this is "much larger than the DeepSeek moment of 2025 because it shows multiple labs can compete with and catch up to well-capitalized model vendors." That trend is real regardless of whether Qwen 3.8's specific claims hold up.
But the trend does not substitute for verified data on any individual model. Whether you pick K3 or wait for Qwen 3.8 to publish its scores, the fact that both options exist changes the economics for everyone building on frontier models.
Beyond the model comparison
If you're not building AI infrastructure and just need a working app, there's a simpler path. Emergent lets you describe an application in plain language and get a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capabilities without managing API keys or comparing token prices.
Skip the model comparisons and API setup. Describe your app and let Emergent handle the rest. Start Building and see how far a prompt gets you.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







