MDL-6265EST.2026 · IDX.594
Language modelIn production

Ling-3.0-flash-VL

InclusionAI (Ant Group) · China · 2026

Vision-and-video version of Ant Group's cheap sparse flagship: 124B parameters, 5.5B active, and a 256K window that the released checkpoint reaches only by stretching a 131K one.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Ling-3.0-flash-VL is the multimodal branch of InclusionAI's current generation, published by the open-source arm of Ant Group, the Alipay company. Weights appeared on Hugging Face on 4 September 2026, quantised fp8, fp4 and int4 builds followed on 8 September, and third-party hosting on OpenRouter started on 10 September. As with the text-only Ling-3.0-flash, the MIT licence exists only as a metadata line in the model card; the repository carries no separate licence file. The backbone is the one readers already know from Ling-3.0-flash: a sparse mixture of experts, 512 routed experts with 8 activated plus one shared, 42 layers alternating Kimi Delta Attention and gated MLA in a 5:1 pattern, hidden size 2560, 32 attention heads, vocabulary 157,184. The maker quotes 124 billion total parameters and 5.5 billion activated per token; the published safetensors index counts 124.8 billion, and the activated figure is slightly higher than the 5.1 billion quoted for the text-only sibling, which is what the added visual stack costs. Vision is handled by a 27-layer ViT with a hidden size of 1152 and 16-pixel patches, merged two-by-two in space and two-by-two in time, then aligned to the language model by a two-layer MLP projector. Temporal order is encoded with VideoRoPE, which is what lets the model answer questions about long clips and locate events inside them rather than merely describe frames. One detail worth recording: in the published configuration the visual tower is declared with the model type qwen3_moe_vit, that is the vision architecture of Alibaba's Qwen3 — a direct competitor of Ant Group in the Chinese model market. The context window deserves care. The model card advertises up to 256K tokens, and the maker's own launch recipe does serve 262,144 — but it gets there by switching on YaRN scaling with a factor of 2.0 over an original window of 131,072, and 131,072 is exactly what the released configuration declares. The text-only Ling-3.0-flash, by contrast, declares 262,144 natively. The split shows up in the wild: on OpenRouter the paid listing of this model reports a context of 131,072 tokens while the free listing reports 262,144. InclusionAI reports that the vision version scores 42 on version 4.1.1 of the Artificial Analysis Intelligence Index against 38 for the text-only model, and presents the gain as evidence that adding sight improves general reasoning; that comparison is the maker's own. Thinking mode is enabled by default and can be turned off per request. Hosting is roughly three times dearer than the text model — about $0.06 per million input tokens and $0.18 per million output against $0.021 and $0.063 — so the vision version is cheap by market standards but no longer the price outlier its text sibling was.

#open weights#MIT license#MoE#multimodal#vision#video#agentic#China
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review