Qwen 3.8 Max vs Kimi K3: What We Can Compare Today and What's Still Missing

Qwen 3.8 Max vs Kimi K3 compared on the evidence available as of July 2026, including specs, pricing, open weights, and the one head-to-head test that exists.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Sakthyapriya Shanmugavadivel
Reviewed by
Sakthy
Published: 
Jul 22, 2026
0
 min read
Table of Contents

This comparison is asymmetric, and that asymmetry is the most important thing to understand before reading further. Kimi K3 has independently verified benchmark scores, published per-token API pricing, and a confirmed open-weight release date. Qwen 3.8 Max has a vendor claim of being second only to Fable 5, a preview endpoint with credit-based pricing, and no published benchmarks of any kind.

Both are multi-trillion parameter Chinese MoE models launched within three days of each other in July 2026. Both target the frontier. Both have committed to releasing open weights, though only K3 has a confirmed date. The evidence available for each is fundamentally different, and this article labels that difference clearly rather than pretending both sides of the comparison sit on equal footing.

TL;DR

  • Kimi K3 scores 57 on the Artificial Analysis Intelligence Index with verified scores across nine evaluations. Qwen 3.8 Max has no published benchmark scores.
  • Alibaba claims Qwen 3.8 is comparable to leading frontier models and second only to Fable 5. That claim shipped with no scores, no named test suite, and no methodology. It has not been independently verified.
  • K3 has standard API pricing: $3/$15 per million tokens. Qwen 3.8 Max is available only through credit-based Token Plan subscriptions ($6-$68/month), with no published per-token rate.
  • The one independent head-to-head test (TrilogyAI's StackPerf architecture benchmark) scored K3 at 83 and Qwen 3.8 at 80.
  • Both promise open weights. K3 has a confirmed date (July 27, Modified MIT). Qwen says "soon" with no date or license.

Qwen 3.8 Max and Kimi K3 at a glance

Spec Kimi K3 Qwen 3.8 Max
Vendor Moonshot AI Alibaba (Qwen team)
Announced July 16, 2026 July 19, 2026
Total parameters 2.8 trillion (MoE) 2.4 trillion (MoE, vendor-reported)
Active per token ~50B (16 of 896 experts) Not disclosed
Context window 1,048,576 tokens ~984K tokens (per integration metadata)
Max output 131K default, up to 1M ~131K tokens (per integration metadata)
Modality Text + images + video Multimodal confirmed (image input tested)
Reasoning control Max only (at launch) low, high, xhigh (default); thinking always on (per Coursiv integration metadata)
API pricing $3.00 / $15.00 per 1M tokens No per-token rate published; credit-based subscriptions only
Cached input $0.30 per 1M Not published
Open weights Modified MIT, confirmed July 27 "Coming soon," no date or license
Intelligence Index 57 (Artificial Analysis, verified) Not scored
Model ID kimi-k3 qwen3.8-max-preview

Specs as of July 2026. Sources: Moonshot K3 blog, Alibaba Qwen announcement, Coursiv integration metadata, TrilogyAI benchmark.

The cells marked "Not scored" and "Not published" reflect genuine gaps in Alibaba's public documentation. Until Alibaba publishes a model card, technical report, or benchmark table, those cells cannot be filled with anything other than speculation.

The benchmark gap: K3 has data, Qwen 3.8 does not (yet)

1. What Kimi K3's independent scores look like

K3's benchmark profile is well-established by this point. It is one of the few open-weight models with a full set of independently verified scores. Artificial Analysis scores it at 57 on the Intelligence Index, fourth overall behind Claude Fable 5 (~60), GPT-5.6 Sol (~59), and roughly on par with Claude Opus 4.8 (~56). Qwen 3.8 Max has no equivalent data. The breakdown across K3's individual evaluations:

Evaluation Kimi K3 Source
Intelligence Index 57 Artificial Analysis (independently verified)
Coding Index 76 Artificial Analysis
Agentic Index 50 Artificial Analysis
GPQA Diamond 94% Artificial Analysis
Terminal-Bench v2.1 85% Artificial Analysis
AA-Briefcase (Elo) 1,543 Artificial Analysis
Humanity's Last Exam 44% Artificial Analysis
AA-Omniscience (non-hallucination) 49% Artificial Analysis

All scores independently verified by Artificial Analysis, July 2026.

K3 also holds #1 on LMArena's Frontend Code Arena (blind developer voting) and leads vendor-reported coding benchmarks from Moonshot's launch table, including 88.3% on Terminal-Bench 2.1 (KimiCode harness), 81.2% on FrontierSWE, and 67.5% on DeepSWE.

For a deeper look at how these scores compare to specific competitors, see our Kimi K3 vs Claude Opus 4.8 and Kimi K3 vs GPT-5.6 Sol, Terra, Luna comparisons.

2. The Qwen3.7-Max baseline (predecessor scores)

With no Qwen 3.8 benchmarks published, the predecessor Qwen3.7-Max provides the most relevant reference point. It launched in May 2026 and has been independently tested.

Evaluation Qwen3.7-Max Source
Intelligence Index (v4.1, current) 46 Artificial Analysis (independently verified)
Intelligence Index (v4.0, at launch) 56.6 Artificial Analysis (prior index version)
GPQA Diamond 92.4% Qwen (vendor-reported)
SWE-bench Verified 80.4% Qwen (vendor-reported)
Terminal-Bench 2.0 69.7% Qwen (vendor-reported, Terminus harness)
Context window 1M tokens Official documentation
Pricing $2.50 / $7.50 per 1M Alibaba Cloud Model Studio

Qwen3.7-Max scores from Artificial Analysis and Qwen's official documentation. The Intelligence Index version matters: Qwen3.7-Max scored 56.6 on v4.0 (May 2026) but 46 on the current v4.1 (July 2026). K3's 57 is measured on v4.1. Use the same index version when comparing.

On the current Intelligence Index v4.1, Qwen3.7-Max sits at 46, which is 11 points behind K3's 57. That is a substantial gap. When the model launched in May 2026, it scored 56.6 on the then-current v4.0, placing it fifth overall. The index methodology changed between versions, which means the two numbers are not directly comparable. What matters for this comparison: on the same index version K3 was tested on, Qwen3.7-Max trails by 11 points, not 0.4.

If Qwen 3.8 truly ranks "second only to Fable 5" as Alibaba claims, it would need to jump roughly 14 points from its predecessor's current v4.1 score. That is a large leap. As AIToolsReview noted, K3 also launched with bold vendor claims that Artificial Analysis later confirmed as roughly accurate (57 vs the vendor-implied near-Fable performance). The same pattern may or may not hold for Qwen 3.8. We do not know yet.

3. The one head-to-head test (TrilogyAI StackPerf)

The only published independent comparison ran both models on the same software architecture benchmark. TrilogyAI gave both models identical frozen repository snapshots (269 files) and a 60-minute window to produce an architectural analysis with evidence citations.

Metric Kimi K3 Qwen 3.8 Max Source
Final score (after factual penalties) 83/100 80/100 TrilogyAI StackPerf
Unsupported claim groups 7 7 TrilogyAI
Repository citations 274 354 TrilogyAI
Tool calls 53 44 TrilogyAI
Failed tool calls 2 (recovered) 0 TrilogyAI
K3 advantage Lifecycle reasoning, revision handling, speed, token efficiency TrilogyAI
Qwen advantage System boundary design, replay metadata, tool discipline TrilogyAI

Results from one matched session per model, July 19, 2026. Source: TrilogyAI.

K3 scored three points higher. It handled revisions, regeneration, and scene history more completely and finished with lower latency and fewer tokens. Qwen defined cleaner system boundaries, captured stronger replay metadata, and had zero failed tool calls.

TrilogyAI's conclusion: their agreement on the core architectural decision was stronger than either report individually, making a case for model diversity on high-value analysis work. This is one test, one task type, one evaluator. It cannot establish a universal ranking, but it does confirm that Qwen 3.8 can complete complex, tool-heavy analysis work at a level close to K3.

Pricing: published rates vs credit subscriptions

This is where a direct comparison becomes impossible. K3 has a standard, published per-token API rate. Qwen 3.8 Max does not. Any cost comparison you see elsewhere that puts a dollar-per-million-token figure next to Qwen 3.8 is either inferring from predecessor pricing or guessing from credit conversion rates that Alibaba has not published.

Kimi K3 Qwen 3.8 Max
Pricing model Standard per-token API Credit-based subscriptions (Token Plan)
Input per 1M $3.00 Not published
Output per 1M $15.00 Not published
Cached input per 1M $0.30 Not published
Subscription plans None required (API key only) Lite $6/mo, Standard $18/mo, Pro $68/mo
Preview discount None (standard pricing from launch) 10% of standard rate during preview; 80% night discount
Cost per task $0.95 (Artificial Analysis) Not measured

K3 pricing from Moonshot. Qwen pricing from Alibaba Token Plan. Qwen credit-to-token conversion is not published.

K3's economics are transparent. You pay $3 per million input tokens and $15 per million output tokens, with cached input at $0.30. You can forecast your bill before running a single request.

Qwen 3.8 Max's economics are opaque. The Token Plan charges monthly subscriptions and dispenses credits, but the credit-to-token conversion rate is not publicly documented. As the eesel AI pricing analysis put it: "cheapest to try, least clear to budget." The 10% preview discount makes experimentation cheap, but production budgeting requires knowing what the standard rate will be, and Alibaba has not published one.

For teams that need to forecast AI costs across models, the Emergent Universal LLM Key handles token routing across Claude, OpenAI GPT, and Google Gemini with transparent per-model pricing.

Architecture and specs: 2.8T vs 2.4T

Both models use Mixture-of-Experts architectures at multi-trillion parameter scale. Beyond that high-level similarity, the comparison is lopsided because K3 has published architecture details and Qwen 3.8 has not. Alibaba has released no technical report, no model card, and no architecture disclosure for Qwen 3.8. Everything below about Qwen's architecture is limited to what Alibaba announced publicly or what third parties extracted from the preview endpoint.

Kimi K3 activates 16 of 896 experts per token (~50B active), uses Kimi Delta Attention (hybrid linear attention in a 3:1 ratio with full attention), Attention Residuals, and Gated MLA for key-value compression. Moonshot published these details in its technical blog and recommends 64+ accelerators for self-hosting. The model footprint is ~1.4 TB before KV cache.

Qwen 3.8 Max reports 2.4 trillion total parameters. That is the extent of confirmed architectural information. Alibaba has not disclosed the active parameter count, expert configuration, attention mechanism, or serving requirements. The model is confirmed multimodal (image input verified by TrilogyAI), and the preview endpoint supports thinking mode with low, high, and xhigh reasoning settings (xhigh is the default, per Coursiv integration metadata). Without a technical report, the architecture cannot be evaluated beyond the parameter headline.

The parameter gap (2.8T vs 2.4T) tells you the checkpoint size, not the inference cost or quality. Active parameters, caching efficiency, and provider infrastructure determine the actual economics and speed of each model.

Open weights, access, and ecosystem

Both models promise open weights. Neither is fully open today. But the gap between "promise" and "confirmed" is wide, and it matters if you are making infrastructure decisions.

Kimi K3 has a confirmed open-weight release on July 27, 2026 under a Modified MIT license (the single attribution clause triggers only above 100 million MAU). Moonshot published the date, the license terms, and the deployment recommendations (64+ accelerators). Until then, K3 is available through the Moonshot API, Kimi.com, Kimi Work, Kimi Code, and OpenRouter. Standard API pricing applies across all channels.

Qwen 3.8 Max has an open-weight commitment with no date, no license, no file location, and no confirmed checkpoint details. Alibaba's X post said "going open-weight soon" without further specifics. It is not clear whether the exact 2.4T Max checkpoint will be the one released, or whether a different variant will ship. The preview is available through Token Plan, Qoder (Alibaba's coding tool), and QoderWork.

The preview endpoint (qwen3.8-max-preview) receives continuous upgrades and will eventually be replaced by a formal model, per Alibaba's Token Plan documentation. That means any test run against the preview today may not reflect the model that eventually ships as the stable release.

For teams looking at the broader landscape of open-weight models, our Kimi K3 alternatives roundup covers six options, and the What is Kimi guide provides context on Moonshot's full model lineup.

When to pick Kimi K3 vs Qwen 3.8 Max

The honest answer for most decisions: pick K3 today, evaluate Qwen 3.8 when its benchmarks and pricing are published.

Your situation Pick Why
Need verified benchmark data to justify the choice Kimi K3 57 Intelligence Index, nine verified evaluations, vendor-reported coding table. Qwen has none.
Need to forecast API costs Kimi K3 Published $3/$15 per-token pricing. Qwen's credit system is not forecastable.
Need open weights with a confirmed date Kimi K3 July 27 under Modified MIT. Qwen says "soon" with no date.
Frontend and UI generation Kimi K3 #1 on LMArena's Frontend Code Arena. Qwen 3.8 has no arena ranking.
Want to experiment cheaply during preview Qwen 3.8 Max 10% preview pricing + 80% night discount. Cheapest way to test a frontier-class model.
Architecture analysis and tool-heavy work Either (test both) TrilogyAI's StackPerf shows both are capable. K3 at 83, Qwen at 80.
Already in the Alibaba/Qwen ecosystem Qwen 3.8 Max Qoder, QoderWork, Token Plan integration. Lower switching cost.
Native video input required Kimi K3 Confirmed text + image + video. Qwen 3.8's video support is not documented.
Betting on the "second only to Fable 5" claim Wait That claim has no published evidence. Wait for independent testing.

Decision matrix based on available evidence as of July 21, 2026. This table will change when Alibaba publishes benchmarks and pricing.

The strategic context matters too. Both models represent the same trend: Chinese labs reaching the frontier with open-weight models at aggressive pricing. As the Emerging Trajectories analysis framed it, this is "much larger than the DeepSeek moment of 2025 because it shows multiple labs can compete with and catch up to well-capitalized model vendors." That trend is real regardless of whether Qwen 3.8's specific claims hold up.

But the trend does not substitute for verified data on any individual model. Whether you pick K3 or wait for Qwen 3.8 to publish its scores, the fact that both options exist changes the economics for everyone building on frontier models.

Beyond the model comparison

If you're not building AI infrastructure and just need a working app, there's a simpler path. Emergent lets you describe an application in plain language and get a production-ready, full-stack product with a real backend, real integrations like Stripe, MongoDB, and Shopify, and code you own. It runs Claude, OpenAI GPT, and Google Gemini under the hood through its Universal LLM Key, so you get frontier model capabilities without managing API keys or comparing token prices.

Skip the model comparisons and API setup. Describe your app and let Emergent handle the rest. Start Building and see how far a prompt gets you.

Was this article helpful?
About the writer
Bhavyadeep Sinh Rathod
Content Manager

SEO Content Manager at Emergent, covering the tools and workflows shaping the next era of vibe coding. 8+ years making complex tech topics discoverable and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

Is Qwen 3.8 Max better than Kimi K3?

There is not enough evidence to answer this definitively. Alibaba claims Qwen 3.8 is second only to Fable 5, which would place it above K3 (Intelligence Index 57). But that claim has no published benchmark scores, no methodology, and no independent verification. The one head-to-head test (TrilogyAI StackPerf) gave K3 a slight edge at 83 vs 80. Until Alibaba publishes real scores, K3 is the only model in this comparison with verified data.

How much does Qwen 3.8 Max cost compared to Kimi K3?

A direct comparison is not possible today. K3 costs $3/$15 per million input/output tokens with standard API access. Qwen 3.8 Max is available through Token Plan subscriptions ($6-$68/month) at 10% of standard pricing during preview, but the standard per-token rate has not been published. As a reference, the predecessor Qwen3.7-Max priced at $2.50/$7.50 per million tokens.

When will Qwen 3.8 open weights be available?

Alibaba has said "soon" without providing a date, license, or file location. K3 has a confirmed release date of July 27, 2026 under a Modified MIT license. Both models will require significant hardware to self-host at their multi-trillion parameter scales.

What is the difference between Qwen 3.8 Max and Qwen3.7-Max?

Qwen 3.8 Max is a newer, larger model (2.4T parameters vs Qwen3.7-Max's undisclosed size). Qwen3.7-Max scored 46 on the current Artificial Analysis Intelligence Index v4.1 (it scored 56.6 on the prior v4.0 at launch). Qwen 3.8 Max has no published scores. Qwen3.7-Max had standard API pricing ($2.50/$7.50); Qwen 3.8 Max is preview-only through Token Plan.

Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql