Z.ai never prints how big its models are — its own weight files do, and the fifth generation is twice the fourth
Published: 8/21/2026 · Source: Z.AI Developer Documentation ↗
Z.ai documents fifteen text models and six vision models on one price page, and describes none of them by size. The model cards list context windows, output limits, modalities and features; the parameter count, the detail every comparison starts from, is simply absent. Press coverage fills the gap with figures the company has never confirmed.
It does not have to stay that way, because Z.ai publishes the weights. The safetensors index of each open model states the total exactly: GLM-4.6 (September 2025) and GLM-4.7 (December 2025) come to 356.8 and 358.3 billion parameters, both built as 92 layers with 160 routed experts plus one shared, eight active per token. The fifth generation doubles that. GLM-5, GLM-5.1 and GLM-5.2 all total roughly 753 billion parameters across 78 layers, with 256 routed experts plus one shared — the same skeleton reused three times, from February to June 2026.
The price table follows the same split. Both fourth-generation models cost $0.60 per million input tokens and $2.20 per million output, and the line goes down to GLM-4.7-FlashX at $0.07/$0.40 and GLM-4.7-Flash, which is free. The fifth generation starts at $1.00/$3.20 for GLM-5 and settles at $1.40/$4.40 for GLM-5.1, GLM-5.2 and GLM-5.3 — 2.3 times the fourth-generation price for 2.1 times the parameters.
One model breaks the pattern, and it is worth noting: GLM-5V-Turbo, the vision-and-coding model from April 2026, has no published weights at all. Z.ai repeatedly says it delivers its results "at a smaller model size" — and that is the one claim in the family which, for now, cannot be checked against a file.