MDL-1702EST.2026 · IDX.423
Language modelIn production

GLM-5.3

Z.ai (Zhipu AI) · China · 2026

Same base model as GLM-5.2, all gains from post-training. The weights arrived on 28 Aug 2026 — not under MIT like the small model, but under a bespoke licence that binds only companies above $10bn.

wujec.ai score

–/10

Community score

no votes yet
Sign in to rate

GLM-5.3 is Z.ai's flagship model, announced on 14 August 2026. Its most unusual feature is what it is not: it is not a new model. Z.ai states plainly that GLM-5.3 runs on the same base as GLM-5.2 and that every improvement comes from post-training. The company's own numbers show what that alone can buy — Terminal-Bench 3.0 rises from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, Agents' Last Exam (CLI) from 23.8 to 28.5. The record Z.ai claims on the first two is expressly a record among open-weight models, not against the field. On the in-house Z.ai Code Bench the model reaches 34.5% at roughly 75,000 output tokens per task against 23.4% at 96,000 for its predecessor: better results while spending fewer tokens. The same chart holds two rows the summary at the top of the announcement leaves out. At high effort GLM-5.3 does pass a closed model — 31.4% at about 50,000 output tokens against 29.5% at 120,000 for Claude Opus 4.8 — but at maximum effort it stays behind Claude Fable 5, which reaches 39.5%. The surprise, by Z.ai's own account, was security. Vulnerability-discovery environments were added to the training mix in the expectation of modest gains; what emerged instead was a model that reasons across whole exploitation chains. On CyberGym it scores 84.5%, up from 77.2%, narrowly ahead of every other model in the company's comparison table (83.8% and 83.6% for the two closed leaders). Deeper into the chain the picture reverses: on ExploitBench GLM-5.3 more than doubles its predecessor to 54.4% but remains far behind the closed frontier (78.0% and 76.5%), and on ExploitGym it completes 105 tasks in two hours and 130 in six, against 29 and 39 for GLM-5.2. Only one closed model is quoted on this last benchmark — Mythos 5, at 181 tasks in two hours and 247 in six — and the budgets are normalised model by model using throughput figures that Z.ai leaves to the footnotes, so the two hours are not the same two hours for each contestant. Z.ai says so itself: capability grows fastest exactly where the gap to the closed frontier is widest. The weights were published on 28 August 2026, two weeks after the announcement, as 141 files totalling 755 GB — 753 billion parameters, released natively in FP8 rather than converted after the fact, with a separate BF16 repository alongside. They do not carry the MIT licence that Z.ai gave the small GLM-5.3-Flash two days earlier. The flagship has its own "GLM-5.3 License": the text is MIT-like in substance — use, modify, distribute, sublicense, sell, fine-tune — with a single condition attached. An operator of a model-as-a-service business whose group revenue exceeds ten billion dollars over any consecutive twelve months must pass a Z.ai security review before any commercial use. Below that threshold the licence imposes nothing. In practice the clause is aimed at a handful of cloud giants and leaves everyone else free. One further caveat matters in migration: GLM-5.3 no longer accepts thinking.type: disabled. Applications that switched thinking off must set a reasoning effort level before moving over, or the request fails.

#open weights#coding#1M context#MoE#cybersecurity
Official website ↗

▸News

▸Videos

No videos yet.

▸Reviews

No reviews yet. Be the first!

Sign in to write a review