MDL-6263EST.2026 · IDX.385
Language modelIn production

Qwen3.8-Max

Alibaba Cloud · China · 2026

Alibaba's largest model to date — a 2.4-trillion-parameter sparse MoE with a one-million-token context window.

wujec.ai score

8.8/10

Community score

no votes yet
Sign in to rate

Qwen3.8-Max, released on 3 August 2026 after a preview at WAIC in Shanghai on 19 July, is the flagship of Alibaba's Qwen line and the largest model the family has produced: 2.4 trillion parameters in a sparse mixture-of-experts arrangement, of which roughly 95 billion are activated per token. It handles up to one million tokens of context with a 131,072-token output limit, and is served worldwide through Alibaba Cloud's Model Studio APIs and the QwenWork agent platform. The pricing is the part worth reading twice: 2.00 USD per million input tokens and 6.00 USD per million output, with cached input reads at 0.25 USD — and that rate is flat across the entire million-token window. Western frontier models charge a premium above a threshold (GPT-5.6 doubles input pricing past 272k tokens, Grok 4.5 past 200k), so the gap between Qwen and its rivals widens the longer the prompt gets. At launch it ranked as the strongest Chinese model for text on the Arena.AI leaderboard and placed second globally on vision tasks. Vendor-reported benchmarks include Terminal-Bench 2.1 at 86.6, PaperBench at 93.0 and IFBench at 82.8. On 9 August 2026 Alibaba uploaded the weights to Hugging Face as Qwen3.8-2.4T-A95B (bf16 and FP8) — the first Max-class Qwen ever released for download — under a bespoke Qwen3.8-Max License rather than a standard open licence; that licence file itself was only added on 12 August. The download is not identical to the rented model: Alibaba's own model card states that the hosted Qwen3.8-Max adds vision input, a non-thinking mode, a million-token context by default and built-in tools, so the published checkpoint is text-only with a native 262,144-token window, and it has its own profile in this catalogue. A compact Qwen3.8-27B variant for hardware-constrained deployments went out separately under plain Apache 2.0. The earlier Qwen 3 line remains available and is documented in its own profile.

#MoE#frontier#long context#multimodal#China
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review