MDL-2834EST.2026 · IDX.997
Language modelIn production

Ling-3.0-flash

InclusionAI (Ant Group) · China · 2026

Ant Group's open-weight hybrid-linear model: 124B parameters, only 5.1B active per token, with a 256K context.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Ling-3.0-flash is the current generation of InclusionAI, the open-source arm of Ant Group — the Alipay company — and the first model of that publisher documented in this catalogue. The weights appeared on Hugging Face on 2 August 2026 under an MIT declaration in the model card; the repository carries no separate licence file, so the terms rest on that metadata line alone. The design point is how little of the model runs at a time. It is a mixture of experts with 124 billion total parameters and 5.1 billion activated per token — by the maker's own accounting about 12.4 per cent of the parameters and 8.1 per cent of the compute of its previous trillion-class flagship, Ring-2.6-1T, which it claims to match or beat on key benchmarks. The published safetensors index actually counts 127.5 billion parameters, slightly above the headline figure. Routing picks 8 of 512 routed experts plus one shared expert. The attention stack is the unusual part. Instead of retro-fitting linear attention after training, InclusionAI pre-trained on it from the start, alternating Kimi Delta Attention and gated multi-head latent attention in a 5:1 pattern — 35 KDA layers to 7 MLA layers. Context was trained in stages, 8K to 32K to 256K, and the released model serves 262,144 tokens. The target is agent work rather than chat. InclusionAI says it trained against more than 10,000 interactive environments for coding, general and deep-research agents, and it ships a hierarchical caching setup (SGLang HiCache with Mooncake) that it credits with cutting time to first token by 60 to over 80 per cent on long inputs. Reported strengths are the SWE-Bench family, Terminal-Bench, MCP-Atlas and other tool-use suites; all published numbers are the maker's own. Thinking mode is on by default and can be switched off per request. Third-party hosting is live: OpenRouter lists the model with a 262K context, and on 27 August 2026 a finance-tuned sibling, Ling-3.0-flash-fin, appeared there as well.

#open weights#MIT license#MoE#reasoning#agentic#256K context
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review