Alibaba gives its rivals the better of two prompts and keeps one for itself — then loses that row to its own previous flagship
Published: 9/13/2026 · Source: Qwen3.8-27B model card, Hugging Face ↗
Benchmark tables published by model makers almost always tilt toward the maker. The card for Qwen3.8-27B, Alibaba's open-weights release of 14 August 2026, contains a footnote that tilts the other way — and the row it governs is one the company loses anyway.
The row is MathVision, a test of visual mathematical problem solving. The footnote reads: Qwen3.8-27B is evaluated using one fixed prompt, while for the remaining models the company reports the higher score from two prompt variants, one requiring a particular answer format and one not. Prompt choice is worth real points on this kind of test, and here the vendor hands that advantage to everyone except itself. The result is 90.0 for Qwen3.8-27B against 90.3 for Qwen3.7-Plus — the closed flagship of the company's own previous generation, printed in bold as the winner of the row.
The same table leans the conventional way a few blocks higher. In the coding section every competing model was re-run by Alibaba itself inside the Claude Code harness, at a fixed temperature and a 256K context window, and the SWE-bench Pro task set was corrected before those runs. One column is exempt: Opus 4.6 Max, which keeps its officially reported 53.4. So the headline coding comparison — 61.7 for Qwen3.8-27B against 53.4 — sets the vendor's own execution against a figure produced by someone else, on a task set the vendor edited. That is disclosed in a footnote too, and it is the kind of asymmetry that usually goes the maker's way.
Two of the table's remaining rows are worth reading for the same reason. Humanity's Last Exam, graded by GPT-4o, gives 30.8 to the new model and 34.7 to Qwen3.7-Plus; GPQA Diamond gives 89.2 against 90.3. The pattern under all three losses is consistent and unsurprising: a 27.78-billion-parameter dense model that anyone can download beats a closed cloud flagship at work that can be executed and graded, and trails it wherever the task is recall.
None of this makes the release smaller than it is. It makes the footnotes load-bearing. Our profiles of Qwen3.8-27B and Qwen3.6-27B were updated today to carry them.