Ling-3.0-tiny
InclusionAI (Ant Group) · China · 2026
The small sibling of Ant Group's Ling-3.0 family: 7.9B parameters, 1.3B active per token, and a footprint that fits a laptop.
Ling-3.0-tiny is the lightweight member of InclusionAI's current family, published on Hugging Face on 10 August 2026, eight days after the larger Ling-3.0-flash. It keeps the same hybrid-linear recipe and shrinks it by roughly a factor of sixteen: 7.9 billion total parameters against 124 billion, and 1.3 billion activated per token against 5.1 billion. The safetensors index confirms the headline figure at 7.89 billion. The point of the model is where it runs. InclusionAI reports validating it on an NVIDIA DGX Spark, an Apple Silicon MacBook and a Mac mini, and quotes roughly 8.34 GiB of peak memory at an 8K context in FP8 — which puts it inside a 16 GB consumer machine rather than a datacentre. On that hardware the maker measures 100 to 105 tokens per second on the DGX Spark and 86 to 90 on an M4 Pro MacBook. Architecturally it alternates three Kimi Delta Attention layers with one gated multi-head latent attention layer in each four-layer block — 24 layers in total, so 18 KDA to 6 MLA. The larger flash model uses a 5:1 ratio. The feed-forward side is a sparse mixture of 128 routed experts, of which each token activates 8, plus one shared expert. The advertised 256K context deserves a footnote. The configuration file in the repository declares a native window of 131,072 tokens; the 262,144 figure quoted in the model card, and used for its own benchmark runs, is reached by switching on YaRN extrapolation with a factor of 2.0 at serving time. Readers running the model unmodified get half the headline number. On the independent Artificial Analysis suite the maker reports a score of 25 on the Intelligence Index v4.1.1 and 16 on the Agentic Index, with output above 160 tokens per second and about 18 seconds end to end for a 500-token answer including thinking time. Reasoning mode is on by default and can be turned off per request. As with the rest of the family, the MIT licence exists only as a metadata line in the model card: the repository contains no separate licence file. In its first month the model was downloaded about 17,300 times, close to the 18,600 of the sixteen-times-larger flash — the small model is being taken up at nearly the same rate.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!