Z.ai credits Claude Opus 4.8 with ten points more than Anthropic does — on the same test, in the same version, and it costs Z.ai the win
Published: 9/13/2026 · Source: Anthropic / Z.ai ↗
Two vendors published a Terminal-Bench 2.1 score for the same model, Claude Opus 4.8, within weeks of each other. Anthropic, in the launch table for that model, gives 74.6%. Z.ai, in the model card for GLM-5.2, writes that its own model "on Terminal-Bench 2.1 (81.0) lands within a few points of Claude Opus 4.8 (85.0)". The gap is 10.4 points, and the two numbers point in opposite directions.
Read through Z.ai's table, GLM-5.2 is the challenger: 81.0 against a rival at 85.0, close but behind. Read through Anthropic's own table, the same 81.0 comfortably beats Opus 4.8's 74.6. In other words, the Chinese lab published a figure for the competitor that makes its own open-weight model look worse than the competitor's own published number would.
Neither table is wrong on its face. Terminal-Bench does not score a model — it scores a model driven by an agent harness, and the official leaderboard says so explicitly: every entry carries an AGENT column next to the model name. Z.ai measured under Terminus-2. Anthropic's launch table does not name a harness at all, and its own row for the same benchmark version hands the top score to GPT-5.5 at 78.2%, ahead of both.
The practical lesson is narrower than "benchmarks are unreliable". It is that a Terminal-Bench figure without a harness attached cannot be compared with another Terminal-Bench figure, even at the identical benchmark version, and even when both come from companies with every reason to measure carefully. On the public leaderboard the same Opus 4.8, run under Claude Code, scores 23.6% — on Terminal-Bench 3.0, a third number for a model that has not changed at all.
We have rewritten both catalogue profiles accordingly: the Opus 4.8 entry now records the row it loses, and the GLM-5.2 entry names the harness behind its 81.0.