MDL-9445EST.2025 · IDX.903
Language modelIn production

LLaDA2.0-mini

InclusionAI (Ant Group) · China · 2025

A 16B open-weight diffusion language model that writes text in 32-token blocks instead of one token at a time — and the most downloaded model its publisher has ever released.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

LLaDA2.0-mini is the small member of the diffusion line from InclusionAI, the open-source arm of Ant Group, and by download count the publisher's single most successful release: over 1.01 million downloads since the weights went up on 25 November 2025, against 7,200 for the company's trillion-parameter reasoning flagship. What sets it apart from almost everything else in this catalogue is how it produces text. Conventional language models are autoregressive: they emit one token, look at it, and emit the next. This one is a diffusion model — it starts from a masked block and unmasks it over repeated refinement passes. The maker's recommended settings generate in blocks of 32 tokens over 32 denoising steps, at temperature 0; raising the temperature, it warns, can make the model mix languages. The weights are a mixture of experts. The published safetensors index counts 16.26 billion parameters in BF16, of which roughly 1.4 billion are active per token — 256 routed experts with 8 selected plus one shared expert, across 20 layers, a hidden size of 2,048 and a 157,184-token vocabulary. Context is 32,768 tokens. That footprint is the practical point: the model was built to run on a single ordinary accelerator rather than a rack. It was not trained from scratch. InclusionAI continued training on its Ling 2.0 base with roughly 20 trillion tokens, then instruction-tuned the result using its own dFactory framework. On the maker's own benchmark table it averages 71.67 against 70.19 for Qwen3-8B without thinking, scoring 80.53 on MMLU, 47.98 on GPQA, 93.22 on MATH and 86.59 on HumanEval; on AIME 2025 it reaches 36.67, well behind the reasoning-tuned models of its own house. All of those figures are the publisher's. The licence is cleaner than at this publisher's other models: the card names Apache 2.0 in prose and metadata, though the repository still ships no separate licence file among its 16 files. The line has since continued — LLaDA2.0-flash is the larger sibling, and LLaDA2.1-mini arrived on 9 February 2026 — yet this older release is still downloaded about twice as often as its successor.

#open weights#diffusion model#MoE#Apache 2.0#on-device#32K context
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review