MDL-6190EST.2025 · IDX.267
Language modelIn production

Ling-mini-2.0

InclusionAI (Ant Group) · China · 2025

The small model that opened Ant Group's Ling 2.0 architecture — 16 billion parameters of which 1.4 billion run per token, and still the publisher's most downloaded conventional model.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Ling-mini-2.0 is the model that introduced the Ling 2.0 architecture on 8 September 2025, and everything InclusionAI has released since stands on it — the trillion-parameter Ling-1T and Ring-1T, the hybrid-attention Ring-2.5-1T and Ring-2.6-1T, and the LLaDA diffusion line, which reuses this model's exact dimensions. In download terms it also remains the publisher's most-used conventional model: 13,058 fetches in the 30 days to 29 August 2026, more than every one of its trillion-parameter siblings put together. The design idea is extreme sparsity. The model holds 16.26 billion parameters but activates 1.43 billion per token, and only 789 million of those outside the embeddings — an activation ratio of one in thirty-two. The maker's claim, derived from its own published scaling laws, is that this buys roughly sevenfold leverage: a model with 1.4 billion active parameters performing like a dense model of 7 to 8 billion. The supporting choices are listed openly — expert granularity, a shared expert, sigmoid routing without an auxiliary loss, multi-token prediction layers, QK normalisation and half RoPE. Training ran on more than 20 trillion tokens, followed by staged supervised fine-tuning and reinforcement learning. Context is 32,768 tokens natively and 131,072 with YaRN. Speed is the practical selling point. On short answers under 2,000 tokens the maker measures over 300 tokens per second on an H20 deployment, more than twice the rate of a dense 8-billion model, and says the advantage grows to sevenfold as sequences lengthen. Training was done in FP8 mixed precision throughout, with loss curves reported as near-identical to BF16 over a trillion tokens, and the FP8 training recipe itself was open-sourced. The release is unusually generous towards researchers rather than users. Alongside the finished model the publisher put out five pre-training checkpoints — the base model before fine-tuning, plus four snapshots taken at 5, 10, 15 and 20 trillion training tokens. Very few makers publish the intermediate stages of a training run, and they are what allow outside researchers to study how a model's abilities appear rather than only what they end up as. The licence follows this publisher's habit. The repository metadata carries an MIT label, but among the 14 files of the weights repository there is no licence document; the licence statement lives in a separate code project on GitHub. Access outside the weights is through the ZenMux marketplace.

#open weights#MoE#small model#FP8#128K context#research checkpoints
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review