All newsResearch

Two models, same day, same licence: the small one was downloaded twenty times more

Published: 9/3/2026 · Source: Hugging Face model repositories and Z.ai price list

Z.ai published its two vision models on the same day, under the same MIT licence, and let anyone download either one. Ten months later the counters say something plain about who open weights are actually for: the small model was pulled about 82,600 times in a thirty-day window, the flagship about 4,000. Nothing about price explains it, because neither download costs anything. The pair is GLM-4.6V, at 107.7 billion parameters, and GLM-4.6V-Flash, at 10.3 billion. Both went out on 7 December 2025. The smaller one is not a stripped-down demo: it carries the same 131,072-token context window and a full 24-layer vision encoder that reads video frames, not just stills. What separates them is the hardware bill. Ten billion parameters fit on a single accelerator, and quantised community builds of this model run on a laptop. A hundred and seven billion need a server. The same shape appears across other publishers. The most-downloaded model in Alibaba's entire Qwen catalogue is the smallest one it ships, Qwen3-0.6B, at about 21.4 million pulls; the rest of the top of the list is 4B, 7B, 8B and 9B. No 235-billion or 480-billion-parameter Qwen appears anywhere in the top sixty — the largest that does is the older Qwen-72B, in twenty-sixth place. At Google, the leader of the Gemma family is gemma-4-31B-it at about 8 million, with the small E2B and E4B variants close behind the mid-range. The line is not drawn where the marketing draws it. Publishers rank their open models by capability and put the biggest at the top; the downloads rank them by whether they fit in the memory of one graphics card. Everything above that threshold is, for most of the people clicking download, a specification rather than a tool. One caveat on the numbers: Hugging Face reports downloads over a rolling thirty-day window, not since publication, so these are current usage rather than lifetime totals — which is precisely why they describe who is running the models today.