MDL-7882EST.2026 · IDX.366
Language modelIn production

Qwen3-VL-Reranker-2B

Alibaba Cloud · China · 2026

The most downloaded of Alibaba's four search models: at two billion parameters it adds almost eight points on scanned-document retrieval over the embedding model of identical size, for one extra stage.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3-VL-Reranker-2B, published on 7 January 2026, is the model that actually runs in production. With 2.07 million Hugging Face downloads in the 30 days to 1 September 2026 it is the most used of Alibaba's four retrieval models — more than the two embedding models and roughly 48 times more than the 8-billion reranker that beats it on every benchmark. Its job is to re-order a shortlist. A first-stage embedding model returns candidates cheaply; this model then reads the query and each candidate together and scores how well they match. Either side may be text, a photograph, a screenshot, a slide or a video clip, which is what distinguishes this suite from a conventional text reranker: a typed question can be scored against a page that exists only as an image. The interesting number is what the second stage buys at this size. Against Qwen3-VL-Embedding-2B — same parameter count, same base model, same family — it gains 7.9 points on ViDoRe v3 (60.8 against 52.9) and 9.9 points on JinaVDR (80.9 against 71.0), both benchmarks of retrieval from scanned and rendered documents. On plain image and video retrieval it is marginally behind, by 1.0 and 1.5 points respectively. Reranking, in other words, pays off precisely where a document is a picture of text — invoices, slides, reports, forms — and barely at all where it is a photograph. The comparison with a direct competitor is more mixed than the vendor's framing suggests, and worth stating: against jina-reranker-m0, also 2 billion parameters, it is ahead on image retrieval (73.8 against 68.2) and on ViDoRe v3 (60.8 against 57.8), but behind on visual document retrieval (83.4 against 85.2) and on JinaVDR (80.9 against 82.2). Structurally it is Qwen3-VL-2B-Instruct with a scoring head: 28 layers, hidden dimension 2,048, a 24-block vision tower, 2.13 billion parameters, supported context of 32,768 tokens, more than 30 languages. It fits comfortably on a single consumer graphics card, which is the whole reason for its popularity. Weights are Apache 2.0 with no European carve-out.

#open weights#reranking#retrieval#multimodal#vision#small model#China
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review