MDL-2168EST.2026 · IDX.637
Language modelIn production

Qwen3.5-2B

Alibaba Cloud · China · 2026

A two-billion-parameter vision-language model under Apache 2.0, small enough for a laptop, that reads documents better than Alibaba's twice-as-large dedicated vision model — and which the vendor itself recommends for prototyping and fine-tuning rather than production.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3.5-2B, published on 28 February 2026, is where Alibaba's own description of its generation changes tone. From this size down the model card adds a sentence absent from every larger sibling: given the parameter scale, the intended uses are prototyping, task-specific fine-tuning, and research or development. The claim that the context window extends to a million tokens is quietly dropped as well; 262,144 tokens natively is all that is promised. That caution is worth reading next to what the model actually does. It is a full vision-language system, and on Alibaba's numbers it beats Qwen3-VL-2B — the previous generation's dedicated vision model of the same size — across essentially every visual test: MMMU 64.2 against 61.4, OCRBench 84.5 against 79.2, CharXiv 58.8 against 37.1, and VlmsAreBlind 75.8 against 50.0. On several document and chart tasks it also passes Qwen3-VL-4B, a model twice its size: OCRBench 84.5 against 80.8, CharXiv 58.8 against 50.3, CountBench 91.4 against 89.4. Where the vendor's warning earns its place is text. With thinking enabled it reports MMLU-Pro 66.5, GPQA 51.6 and IFEval 78.6 — respectable for the size, but well behind the older Qwen3-4B-2507 on each. Competition mathematics is where the floor shows: HMMT February 2025 at 22.9 against 57.5. This is a model that reads and describes reliably and reasons only in short steps. The weight files hold 2.27 billion parameters in bfloat16, including the vision encoder — roughly 4.5 GB on disk, so a mid-range laptop graphics card or an Apple machine with unified memory will run it without quantisation. Structurally it follows the generation: 24 layers as six repetitions of three Gated DeltaNet blocks and one gated-attention block, hidden dimension 2,048, eight query heads to two key-value heads, vocabulary 248,320, output embeddings tied to the input side. The licence is Apache 2.0 and the licence file in the repository is the unmodified Apache text.

#open weights#dense#multimodal#local deployment#China#small model
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review