MDL-2601EST.2025 · IDX.854
Language modelIn production

Qwen3-VL-8B-Instruct

Alibaba Cloud · China · 2025

The most downloaded model Alibaba has ever published: an 8.8-billion-parameter vision-language model that reads documents in 32 languages, operates phone and desktop interfaces, and runs on a single consumer graphics card.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3-VL-8B-Instruct, published on 11 October 2025, is the workhorse of Alibaba's vision-language line and, by download count, the single most popular model the company has ever released. In the thirty days to 2 September 2026 it was fetched 9.58 million times from Hugging Face — more than any Qwen chat model of any generation, and more than the entire Granite family of IBM put together. The reasons are practical rather than glamorous. It carries 8.77 billion parameters, all active, across 36 transformer layers with a hidden dimension of 4,096, which puts it inside the memory budget of a single 24 GB consumer graphics card in bf16 and comfortably inside a laptop in quantised form. Its context window is 262,144 tokens natively, which the vendor says can be stretched to a million. And it does the unglamorous jobs that businesses actually pay for: optical character recognition across 32 languages, up from 19 in the previous generation, with the vendor claiming robustness in low light, blur and tilt, and better handling of rare and archaic characters. The more ambitious claims concern agency. Alibaba positions the model as a visual agent that recognises interface elements on PC and mobile screens, understands what they do and drives them to complete a task; it also generates Draw.io diagrams, HTML, CSS and JavaScript from a screenshot or a video frame. Spatial work is a stated priority — 2D grounding is described as stronger, and 3D grounding is offered for robotics and embodied applications. For video, a text-timestamp alignment scheme replaces the earlier positional approach, allowing the model to locate an event at second-level resolution in footage hours long. One architectural detail is worth recording, because the vendor does not draw attention to it. The configuration files of this model fuse visual features from three intermediate layers of the vision tower — the mechanism Alibaba calls DeepStack — into the language model. In every flagship the company has shipped since February 2026, from Qwen3.5 onwards, vision is built into the main model but that fusion list is empty. The separate vision line was folded into the flagship, and the fine-grained fusion it advertised did not travel with it. Whether that costs anything in practice is not something Alibaba has published numbers on; the download figures suggest a great many users have not moved on. Benchmark results for this model are published by the vendor only as chart images rather than tables, so this profile does not quote scores. The weights are Apache 2.0, with no European carve-out and no additional restriction.

#open weights#dense#multimodal#vision#local deployment#China
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review