MDL-8258EST.2025 · IDX.346
Language modelIn production

Ling-1T

InclusionAI (Ant Group) · China · 2025

The first trillion-parameter open-weight model from Ant Group's research arm — and the one that answers without thinking out loud first, by design rather than by omission.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Ling-1T, published on 2 October 2025, was the first model at the trillion-parameter mark from InclusionAI, the open-source arm of Ant Group, and it opened the line that the catalogue's Ring-2.6-1T and Ling-3.0 models later continued. It is deliberately a non-thinking model. Where most flagship releases of late 2025 competed by generating longer and longer chains of reasoning before answering, this one was tuned to answer directly, and its maker frames that as the point: it claims to extend the trade-off curve between accuracy and reasoning length rather than to climb it. The training method behind that is called Evo-CoT, an escalating chain-of-thought curriculum applied in mid- and post-training, followed by reinforcement learning at the level of whole sentences rather than tokens or sequences — a method the team named LPO. The weights are a mixture of experts with a 1-in-32 activation ratio: about 50 billion parameters do the work on any given token. The published safetensors index counts 999.71 billion parameters, so the round trillion in the name is a ceiling rather than a measurement. Pre-training ran on more than 20 trillion tokens, over 40 per cent of them reasoning-heavy in the later stages. Context is 32,768 tokens natively and 131,072 with YaRN extension. The maker states this is the largest foundation model trained in FP8 precision to date, which it credits with a 15 per cent end-to-end speed-up and a loss deviation under 0.1 per cent against BF16. One thing a reader should know before looking at the model card: it publishes no benchmark table. Every result is a chart image — and the card says those charts were drawn by Ling-1T itself, as a demonstration of its front-end code generation. The only figures given in prose are a roughly 70 per cent tool-call accuracy on BFCL v3, reached with light instruction tuning and no large-scale trajectory data, and a first place among open-source models on ArtifactsBench. Both are the publisher's own claims. The licence follows this publisher's usual pattern, which is worth stating plainly: the metadata label reads MIT, the card text has no licence section at all, and the repository ships no licence file among its 165 files. Running the model needs a cluster — the maker's own instructions assume four nodes of eight accelerators. Downloads are correspondingly modest, about 2,500 in the 30 days to 29 August 2026, against 544 likes on the repository.

#open weights#MoE#trillion parameters#MIT#128K context#FP8
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review