Research9/13/2026 · Z.ai — GLM-5.3 model guideZ.ai measures its model against a rival in two hours — but the two hours are set separately for each model
The most honest sentence in Z.ai's GLM-5.3 announcement is also the one that makes its own headline number hard to read: on the ExploitGym benchmark, the time budget every model gets is normalised model by model, using throughput figures the company leaves to the footnotes. The comparison reads as "105 tasks in two hours against 181" — but the two hours are not the same two hours for each contestant.
ExploitGym counts how many exploitation tasks a model completes under a fixed time budget. Z.ai reports GLM-5.3 at 105 tasks in two hours and 130 in six, up from 29 and 39 for GLM-5.2. Only one closed model appears in that row: Mythos 5, at 181 and 247. A benchmark scored in wall-clock time measures two things at once — how well a model reasons and how fast it is served — and normalising the budget is a defensible way to separate them. It also means the headline gap depends on a throughput assumption the reader never sees.
Two more qualifications sit in the same announcement, both of them stated by Z.ai and both routinely dropped when the numbers travel. The state-of-the-art claim on Terminal-Bench 3.0 and Agents' Last Exam is expressly a record among open-weight models, not against the field. And on the company's in-house Z.ai Code Bench, GLM-5.3 reaches 34.5% at maximum effort — ahead of Claude Opus 4.8 at the high setting, behind Claude Fable 5 at 39.5%. The summary at the top of the page mentions the first comparison and not the second.
None of this makes the security result small. GLM-5.3 climbs from 77.2% to 84.5% on CyberGym, the best figure in the company's table, and Z.ai says the model found 2,436 vulnerabilities across 269 real projects, 1,097 of them medium-to-high severity. The company's own reading is the useful one: capability grows fastest exactly where the distance to the closed frontier is greatest. We have corrected our GLM-5.3 profile accordingly — the earlier version attributed the ExploitGym figures 181 and 247 to two different closed models, when both belong to one, at two different budgets.
GLM-5.3 →