Qwen3-VL-2B-Instruct
Alibaba Cloud · China · 2025
The smallest model in Alibaba's vision line, at 2.1 billion parameters: small enough for a phone, and the only variant whose eye was shrunk along with its brain.
Qwen3-VL-2B-Instruct, published on 19 October 2025, is the bottom of Alibaba's vision-language range and the version most likely to end up on a device rather than in a data centre. It carries 2.13 billion parameters across 28 layers with a hidden dimension of 2,048 — a footprint that fits in a few gigabytes once quantised, which is why the vendor also ships a GGUF build for laptop and phone runtimes. The context window is the same 262,144 tokens as its far larger siblings. One difference from the rest of the family is visible only in the configuration files and deserves stating plainly, because it is the kind of thing a specification sheet hides. In the 8-billion and 235-billion versions the vision tower is 27 blocks deep and fuses features into the language model from layers 8, 16 and 24. Here the tower is 24 blocks deep and fuses from layers 5, 11 and 17. The small model does not merely think with fewer parameters; it also looks with a smaller eye. Alibaba does not quantify what that costs, and this profile does not guess — but a reader comparing variants should know that the reduction is on both sides, not just the language side. The feature list is otherwise the generation's: optical character recognition across 32 languages, operation of desktop and mobile interfaces as a visual agent, generation of Draw.io diagrams and front-end code from a screenshot, spatial grounding in two and three dimensions, and second-level event location in long video through text-timestamp alignment. Adoption is substantial for a model of this size: 2.93 million Hugging Face downloads in the thirty days to 2 September 2026, roughly a third of the 8-billion version's traffic and more than triple that of the line's own flagship. Weights are Apache 2.0, with no European carve-out. Benchmarks are published by the vendor only as chart images, so no scores are quoted here.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!