One company, five launches, four different ways of measuring the competition — all of them disclosed, none of them in the headline
Published: 9/13/2026 · Source: Mistral AI — launch posts for Pixtral Large, Devstral, Mistral Medium 3, Devstral 2 and Mistral Small 4 ↗
We spent today re-checking benchmark figures in our Mistral profiles against the company's own launch posts, one number at a time. The numbers held up. What did not hold up was the idea that they can be read side by side. Across five announcements the same company measured its rivals four different ways, and each time it said so — a sentence or two below the headline that everyone quotes.
Pixtral Large, November 2024, is the strict case: the rivals were run "through a common testing harness", and the 69.4 percent on MathVista, plus the wins over GPT-4o and Gemini 1.5 Pro on ChartQA and DocVQA, mean what a reader assumes they mean. Devstral, May 2025, is the loose case. Devstral's own 46.8 percent on SWE-bench Verified was measured in the OpenHands scaffold; the table that puts it ahead of closed models compares, in Mistral's own words, models "evaluated under any scaffold (including ones custom for the model)". On this benchmark the scaffold — the harness that feeds the model the repository and runs the tests — is worth several points on its own, which is why the comparison is not like-for-like, and why the announcement says so.
Mistral Medium 3, May 2025, is the mixed case, and it is the one where the small print argues with itself. The text explains that the company used "numbers reported previously by other providers wherever available" and its own harness for the rest. The asterisk under the chart on the same page reads: "Performance accuracy on all benchmarks were obtained through the same internal evaluation pipeline." Both sentences cannot be true of the same table. The headline claim — at or above 90 percent of Claude Sonnet 3.7 across benchmarks — rests on whichever one is.
Mistral Small 4, this year, adds the fourth variant: a mode asymmetry. The chart showing 0.72 on AA LCR at 1.6 thousand characters of output is labelled as Mistral Small 4 with reasoning, which the model does not do by default — with reasoning effort set to none it behaves, Mistral says, like Mistral Small 3.2. What setting the rival models ran under, the announcement does not say.
None of this is misconduct, and the point is not that Mistral is unusual; it is that a benchmark number is a measurement plus a method, and the method travels badly. It survives the launch post and dies in the summary. The counter-example comes from Mistral too: in the Devstral 2 launch of December 2025 the company commissioned human evaluations through an independent annotation provider, with every model scaffolded through Cline — one rule for everyone — and published the result, which was that its model beat DeepSeek V3.2 and that Claude Sonnet 4.5 remained significantly preferred. A common harness is what makes a loss publishable. Six of the eight Mistral profiles in this catalogue were edited today to carry the method next to the number.