GLM-5.2 lost seven points on a coding benchmark without a single weight changing. The benchmark grades on a curve
Published: 9/13/2026 · Source: FrontierSWE V1 leaderboard (Proximal); Z.ai — GLM-5.2 model card and GLM-5.3 launch post ↗
Z.ai publishes the same benchmark twice, with two different numbers. The GLM-5.2 model card, released with the model in June, gives FrontierSWE 74.4, against 75.1 for Claude Opus 4.8 and 72.6 for GPT-5.5 — Z.ai's own text at the time said the model trailed Opus 4.8 by about one point. The comparison table in the GLM-5.3 launch post, published two months later, gives GLM-5.2 67.5 on the same benchmark, against 66.5 for Opus 4.8. The model did not change. The order did: GLM-5.2 is now the one in front.
The explanation is in the metric, and FrontierSWE states it plainly on its leaderboard. The headline number is not a pass rate. It is dominance — the win rate against a random opponent on a random task, across 17 tasks. A relative score of that kind is rewritten for everybody each time a competitor is added to the field. Between June and August the board gained GLM-5.3, Grok 4.6, Kimi K2.6 and others; GLM-5.2 fell from 74.4 to 67, Opus 4.8 fell from 75.1 to 67, and because Opus fell further, the two swapped places. Neither result is wrong and neither company misquoted anything. The footnote on Z.ai's card is honest about it: the score is given as of 16 June 2026.
The rest of the current V1 board reads the same way, in whole percentages: Claude Fable 5 88, GLM-5.3 and Grok 4.6 78, Grok 4.5 72, then GLM-5.2 and Opus 4.8 at 67 and GPT-5.5 at 65. It is worth knowing what these percentages sit on. On the five implementation tasks in V1, no model completed any task in any trial, so the ranking there falls back to the share of tests passed by the best of five attempts.
FrontierSWE has since moved to V2, and moved away from the curve: 34 tasks, five trials each, a twenty-hour budget per trial, and an absolute score. On that board Claude Fable 5.1 leads with 56.3 percent, give or take 11.1, at an average of 138.55 dollars and 11.6 hours per task, with GPT-5.6 second at 32.2. Same name, three different things a reader could mean by it. Our GLM-5.2 profile now carries both figures with the date and the metric attached, which is the only form in which a relative score means anything.