Alibaba's most downloaded model belongs to a line the company has stopped making
Published: 9/2/2026 · Source: Alibaba, karty i pliki konfiguracyjne modeli Qwen3-VL, Qwen3.5 i Qwen3.8 (Hugging Face) ↗
The most downloaded model Alibaba has ever published is not a chatbot and does not belong to the company's current generation. It is Qwen3-VL-8B-Instruct, a vision-language model from October 2025, fetched 9.58 million times from Hugging Face in the thirty days to 2 September 2026 — more than any Qwen chat model of any generation, and more than the whole Granite family of IBM combined.
That would be unremarkable if the line were still being made. It is not. In February 2026 Alibaba released Qwen3.5, a flagship line in which vision is no longer a separate product: the main models see images and video natively. Since then the company has published no new VL flagship. The last one, Qwen3-VL-235B-A22B, dates from September 2025.
The interesting part is what did not make the move. The VL line was built around a mechanism Alibaba calls DeepStack, which fuses visual features from three intermediate layers of the vision tower into the language model — the vendor presents it as the source of fine-grained image-text alignment. Open the configuration files of the models that replaced the line and the fusion list is empty: Qwen3.5-9B, Qwen3.6-27B and Qwen3.8-27B all carry a vision tower of the same depth, and all of them fuse from nowhere. The mainline models see, but not the way the vision line saw.
Alibaba has published no comparison between the two approaches and no explanation of the change, and this is not a claim that the newer models are worse at looking at pictures — we have not measured that and the vendor's own benchmark charts do not address it. What can be said is narrower and still worth saying: a design the company advertised as its visual advantage was dropped without comment, and the users voting with their downloads have largely stayed with the older architecture.
There is a second detail for anyone choosing a size. The 2-billion version of the vision line, published on 19 October 2025, does not merely have a smaller language model. Its vision tower is 24 blocks deep against 27 in the 8-billion and 235-billion versions, and fuses from layers 5, 11 and 17 rather than 8, 16 and 24. The small model looks with a smaller eye, which no specification sheet mentions.
wujec.ai has today added profiles of all three: the 8-billion workhorse, the 235-billion flagship and the 2-billion version for phones and laptops. All are Apache 2.0, with no European carve-out. Benchmark results are published by Alibaba only as chart images rather than tables, so our profiles quote no scores.