A 27B model beats Alibaba's own 397B flagship — and ties Claude 4.5 Opus on two agent tests
Published: 9/2/2026 · Source: Alibaba, karta modelu Qwen3.6-27B (Hugging Face) i komunikat qwen.ai ↗
Qwen3.6-27B is a dense model with 27.78 billion parameters, all of them working on every token. In Alibaba's own benchmark tables it beats Qwen3.5-397B-A17B, the company's previous-generation open-source flagship with fifteen times more parameters in store, on every major coding benchmark. Both models are measured in the same table, by the same team, with the same agent scaffold — which makes the comparison unusually clean.
The numbers: SWE-bench Verified 77.2 against 76.2, SWE-bench Pro 53.5 against 50.9, SWE-bench Multilingual 71.3 against 69.3, Terminal-Bench 2.0 59.3 against 52.5, NL2Repo 36.2 against 32.2. On SkillsBench the gap is not a gap but a chasm: 48.2 against 30.0. On the vendor's internal front-end generation rating, 1487 against 1186.
Two of those figures deserve separate mention, because the comparison there is not with an open model. Against Claude 4.5 Opus, Terminal-Bench 2.0 ends in a tie at 59.3, and SkillsBench goes to the Chinese model, 48.2 against 45.3.
What this does not show is a 27B model catching up in general. Read the knowledge rows and the picture reverses: MMLU-Pro 86.2 against Opus's 89.5, Humanity's Last Exam 24.0 against 30.8, SuperGPQA 66.0 against 70.6. The pattern is consistent and worth naming. Narrow, verifiable, tool-driven work — patching a repository, driving a terminal — is exactly what reinforcement learning can grade automatically, and it is there that parameter count has stopped deciding the outcome. Broad recall of facts still costs parameters, and the gap there has not closed.
One longer line is visible across this year at a single size. Alibaba shipped a 27-billion-parameter model three times in six months: February, April and August. On SWE-bench Pro the score went 51.2, then 53.5, then 61.7 — the same amount of hardware, a ten-point rise. The caveat is that the August figures come from a different card and partly different benchmark versions, so only the February-to-April step is measured on one table.
Weights are Apache 2.0. All figures above are vendor-reported and have not been independently reproduced.