MDL-6625EST.2025 · IDX.747
Language modelIn production

Qwen3-VL-235B-A22B-Instruct

Alibaba Cloud · China · 2025

The flagship of Alibaba's separate vision line: 235 billion parameters of which 22 billion work on each token, released with open weights five months before the company folded vision into its main models.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3-VL-235B-A22B-Instruct, published on 22 September 2025, opened Alibaba's third vision-language generation and remains the largest model the company has released specifically for seeing rather than talking. It is a mixture of experts: 235.67 billion parameters in total across 94 layers and 128 experts, of which eight are consulted per token, so roughly 22 billion are actually at work on any given step. The context window is 262,144 tokens natively, which the vendor says can be pushed to a million. What the generation introduced is easier to describe than to measure. Interleaved MRoPE distributes positional information across time, width and height instead of privileging one axis, aimed at reasoning over long video. DeepStack fuses features from three intermediate layers of the vision tower — the eighth, sixteenth and twenty-fourth — into the language model, which the vendor presents as the source of its fine-grained image-text alignment. A text-timestamp alignment scheme replaces the earlier approach for locating events in footage, and optical character recognition was extended from 19 languages to 32. The strategic reading matters more than any single number. Alibaba shipped this model in September 2025, followed it with dense variants down to two billion parameters in October, and then in February 2026 released Qwen3.5 — a flagship line in which vision is no longer a separate product but a component of the main model. From that point the company has not published a new VL flagship. Read the configuration files of the models that replaced it and the DeepStack fusion list is empty: the mainline models see, but they do not use the multi-level fusion this line was built around. For a reader choosing today, that leaves an awkward but honest picture. This model is the last of a line, no longer receiving successors, but still the only open-weights model from the vendor built from the ground up for visual work at frontier scale. In the thirty days to 2 September 2026 it was downloaded 825,000 times — an order of magnitude below its own 8-billion sibling, which is what one expects of a model that needs a server rack rather than a graphics card. Benchmark results are published by the vendor only as chart images rather than tables, so this profile quotes no scores. Weights are Apache 2.0, with no European carve-out.

#open weights#mixture of experts#multimodal#vision#China
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review