MDL-3487EST.2026 · IDX.039
Language modelIn production

Qwen3.5-0.8B

Alibaba Cloud · China · 2026

The smallest model of the Qwen3.5 generation: under a billion parameters, about 1.7 GB of weights, and still able to read a scanned document nearly as well as a dedicated vision model twice its size — while its mathematics collapses to nothing.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3.5-0.8B, published on 28 February 2026, is the floor of Alibaba's current generation and the clearest illustration of what survives when a model is shrunk that far. The weight files hold 873 million parameters in bfloat16, including the vision encoder — about 1.7 GB, small enough for a phone-class device, an embedded board or a browser runtime. What survives is perception. On the vendor's numbers the model scores 74.5 on OCRBench and 79.1 in its non-thinking mode, against 79.2 for Qwen3-VL-2B, a dedicated vision model with more than twice the parameters. Document parsing holds up in the same way (OmniDocBench 61.0 thinking, 70.6 non-thinking), and so does counting objects in a photograph (CountBench 77.0). For a reader who wants a local model to pull numbers off a scan or label images in bulk, this is the cheapest thing in the catalogue that does it. What does not survive is reasoning. GPQA falls to 11.9, PolyMATH to 8.2, and on the HMMT competition-mathematics sets Alibaba reports no score at all — the dashes in its own table are the honest answer. Instruction following lands at IFEval 44.0, so even obedience to a complicated prompt is unreliable. The model card carries the vendor's own caveat, shared with the 2B and absent from everything larger: at this parameter scale the intended uses are prototyping, task-specific fine-tuning, and research or development. The million-token context extension claimed for the 4B and above is not claimed here either; 262,144 tokens natively is the promise. The architecture is the generation's hybrid in miniature: 24 layers as six repetitions of three Gated DeltaNet blocks and one gated-attention block, hidden dimension 1,024, feed-forward width 3,584, vocabulary 248,320 — a vocabulary table that, at this size, accounts for a large share of the parameter count. Weights are Apache 2.0, with the unmodified licence text in the repository.

#open weights#dense#multimodal#local deployment#China#small model
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review