GLM-5.1
Z.ai (Zhipu AI) · China · 2026
The GLM version built to keep working: Z.ai says it can run a single task autonomously for up to eight hours.
GLM-5.1 arrived in April 2026 as the third step of Z.ai's fifth generation, and it is the one where the company changed the question it was answering. Instead of chasing single-turn cleverness, the model is tuned for long-horizon work: planning, executing, testing, fixing and delivering, without a human restarting it every few minutes. Z.ai claims it can run autonomously on one task for up to eight hours and describes it as the first Chinese model to reach that level of sustained execution. The company also positions its general and coding ability as aligned overall with Claude Opus 4.6, and reports 58.4 on SWE-Bench Pro — figures worth reading as vendor claims, since they come from Z.ai's own documentation. The published specification is deliberately plain: text in, text out, a 200,000-token context window, up to 128,000 tokens of output, thinking modes, streaming, function calling, context caching, structured JSON output and MCP tools. Pricing is $1.40 per million input tokens and $4.40 per million output, with cached input at $0.26 — the same tier the company later kept for GLM-5.2 and GLM-5.3. What the documentation never mentions is the size of the model, and here the weights answer for it. Z.ai published GLM-5.1 on Hugging Face under an MIT licence: the safetensors index totals 753.9 billion parameters, and the configuration file describes a mixture-of-experts network of 78 layers with 256 routed experts plus one shared, eight active per token, a 154,880-token vocabulary and a sparse-attention architecture. Structurally it is the same skeleton as GLM-5 — the gain is in training and in how long the model can hold a goal, not in a bigger network.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!