The chart that made Opus 5 state of the art has a footnote: on refusals, the answer came from Opus 4.8
Published: 9/13/2026 · Source: Anthropic ↗
The headline claim of Anthropic's Claude Opus 5 launch is a software-engineering chart: on Frontier-Bench v0.1, the company writes, Opus 5 "surpasses all other models, and more than doubles Opus 4.8's performance at a lower cost per task". At the bottom of the same page, in a single footnote under the heading Footnotes, Anthropic explains how that run was conducted — and one sentence changes how the chart should be read.
The footnote says: "These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5."
Read plainly, that means some of the work behind the Opus 5 curve was not done by Opus 5. When a safety classifier blocked a request, the task was handed to the previous flagship, Opus 4.8, and the attempt continued. The same arrangement applied to Claude Fable 5. Anthropic does not say how often it triggered, on which tasks, or what the chart would look like without it — and since the score is a mean reward over five attempts per task, a substitution does not have to be frequent to move a line.
The awkward part is the comparison itself. The number Anthropic advertises is a doubling of Opus 4.8's performance, and Opus 4.8 is the model that stepped in whenever Opus 5 was refused. The footnote also names only the two Anthropic models. What happened when a competing model in the same chart hit its own refusal is not stated; the natural reading is that it simply scored a failed attempt, which would make the two Anthropic entries the only ones in the chart with a second route to an answer.
None of this is concealed, and none of it is an accusation of fabrication. It is disclosed, in the vendor's own words, in the place where methodology belongs. It is also worth noticing that the same release note introduces automatic fallbacks as a product feature: developers can now have requests flagged by safety classifiers on Opus 5 or Fable 5 route automatically to another model rather than being blocked. Anthropic is, in effect, benchmarking the model the way it expects the model to be deployed — as part of a routed system rather than alone. That is a defensible choice, and it is a different thing from what a bar on a chart labelled with one model's name usually means.
Two further qualifications sit in the same footnote and are easy to skip. The run is internal, not a submission to the public Frontier-Bench leaderboard. And it uses the mini-SWE-agent harness — the same variable that, in a case we described yesterday, produced three different Terminal-Bench numbers for one unchanged model.
We have rewritten the catalogue entry for Claude Opus 5 accordingly. It now records the fallback arrangement behind the Frontier-Bench figure, the two rows Anthropic's own announcement does not win — cybersecurity, where Opus 5 stays behind Claude Mythos 5, and CursorBench 3.2, where it lands just under Claude Fable 5 — and the fact that its widely quoted Artificial Analysis Intelligence Index score of 61 belongs to version 4.1 of that index. Under the current version 4.3, the same configuration scores 51.