Qwen3-VL-Embedding-8B
Alibaba Cloud · China · 2026
Alibaba's largest search-indexing model: it turns text, screenshots, photographs and video into one shared set of coordinates — but it is weaker at pure text than the company's own text-only model of identical size.
Qwen3-VL-Embedding-8B, published on 7 January 2026, does not hold a conversation. It converts a document into a list of numbers — a vector — so that a search engine can find related material by measuring distance between vectors. What makes this version unusual is that the document does not have to be text: a photograph, a screenshot, a scanned invoice, a slide or a video clip lands in the same coordinate space as a sentence, which means a text query can retrieve an image and an image query can retrieve a paragraph. The model is built on Qwen3-VL-8B-Instruct and inherits its structure: 36 layers, hidden dimension 4,096, a 27-block vision tower fusing features into the language model from layers 8, 16 and 24. It carries 8.14 billion parameters — around 622 million fewer than its reranking twin, because a model that only produces vectors has no use for the output layer that turns hidden states back into vocabulary tokens, and Alibaba simply left it out of the weights file. The vendor states a working context of 32,768 tokens; the configuration files permit far more, but 32K is the figure the producer supports. Output vectors are up to 4,096 numbers long and the length is adjustable down to 64, so a project short on memory can trade precision for index size without retraining anything. On the vendor's own multimodal retrieval benchmark (MMEB-V2, 78 datasets) it scores 77.9 overall, the highest figure in the comparison table Alibaba published — ahead of the proprietary Seed-1.6-embedding at 76.9 and of every open competitor listed. One result in that same card deserves a reader's attention, because it argues against using this model. On MMTEB, the multilingual text benchmark, it scores 67.88, while Alibaba's own text-only Qwen3-Embedding-8B — same company, same parameter count, published earlier — scores 70.58. The eyes are not free. If a search index contains nothing but text, the older text model is the better instrument, and the vendor's numbers say so plainly. The model supports more than 30 languages and accepts task-specific instructions, which the vendor says are worth 1 to 5 percent accuracy and should be written in English even for other-language material. Weights are Apache 2.0 with no European carve-out. Adoption in the 30 days to 1 September 2026: 1.22 million Hugging Face downloads and 475 likes.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!