Qwen3.5-122B-A10B
Alibaba Cloud · China · 2026
The second-largest model of Alibaba's Qwen3.5 generation: 125 billion parameters of which ten work per token, leading everything below the flagship on knowledge, reasoning and visual mathematics.
Qwen3.5-122B-A10B, published on 24 February 2026, sits one rung below the generation’s flagship, the 397B-A17B, which Alibaba publishes openly as well. The weight files hold 125.1 billion parameters; each token is routed to eight of 256 experts plus one shared expert, so about ten billion are active at any moment. Wśród modeli mniejszych od flagowca prowadzi tam, gdzie najbardziej liczy się pojemność. Alibaba reports MMLU-Pro 86.7, SuperGPQA 67.1, GPQA Diamond 86.6, Humanity's Last Exam with chain of thought 25.3 and with tools 47.5, and Terminal Bench 2 at 49.4 — the last of these well ahead of the hosted GPT-5-mini at 31.9 and of GPT-OSS-120B at 18.7. On agentic work it posts BFCL-V4 72.2 and Browsecomp 63.8. Where it does not lead is instruction adherence and mathematics under pressure, both of which the smaller dense 27B takes (IFEval 95.0 against 93.4, PolyMATH 71.2 against 68.9) — a reminder that in this generation size is not a single ranking. Vision is trained in rather than bolted on, and the results are the strongest in the family: MMMU 83.9, MMMU-Pro 76.9, MathVision 86.2, against 80.6, 69.3 and 74.6 for Qwen3-VL-235B-A22B, the dedicated vision model of the previous generation that carries nearly twice its parameters. Structurally it is 48 layers as twelve repetitions of three Gated DeltaNet blocks and one gated-attention block, each followed by the expert mixture; hidden dimension 3,072, expert width 1,024, 32 query heads to two key-value heads, vocabulary 248,320, trained with multi-token prediction. Context is 262,144 tokens natively, extensible to 1,010,000 per vendor. The licence is Apache 2.0 with the unmodified licence text in the repository. Downloads run at roughly 700,000 in a 30-day window — a fraction of the small models, which is what one expects of a model that needs server-class memory rather than a desktop.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!