All newsResearch

The new open transcriber beats Whisper — until you leave the big languages

Published: 9/5/2026 · Source: Qwen3-ASR model card, Hugging Face

Alibaba's Qwen3-ASR is being read as the model that finally unseats Whisper-large-v3, the OpenAI transcriber that has been the open default since 2023. On the headline numbers it does: 4.90 per cent word error rate against 5.27 on the Fleurs benchmark. But the producer's own results table contains three lines called Fleurs, Fleurs† and Fleurs††, and they are not three tests of the same thing — they are the same test run over 12, 20 and 30 languages. Which line you quote decides who wins. Over the twelve largest languages — English, Chinese, Cantonese, Arabic, German, Spanish, French, Italian, Japanese, Korean, Portuguese, Russian — the new model wins, 4.90 to 5.27. Add eight more, including Polish, Dutch, Hindi, Turkish and Vietnamese, and the win shrinks to a statistical draw: 6.62 to 6.85. Add the last ten — Czech, Danish, Greek, Persian, Finnish, Filipino, Hungarian, Macedonian, Romanian, Swedish — and the order reverses hard: 12.60 against Whisper's 8.16. The averages let the last group be estimated. If each language weighs the same, the ten smallest work out to roughly 25 per cent errors for the new model against about 11 for the three-year-old one — one word in four against one in nine. That is the difference between a transcript a person edits and a transcript a person retypes. None of this makes the release weaker than it looks in the other direction: the same model handles streaming and offline transcription, covers 22 Chinese dialects, transcribes singing and songs over backing music, identifies the spoken language more reliably than Whisper (97.9 against 94.1 per cent), and ships under Apache 2.0. The point is narrower and it is the one a reader actually needs: a state-of-the-art claim in speech recognition is only as broad as the language list underneath it, and that list is usually a footnote. Anyone working in Czech, Hungarian, Romanian or the Nordic languages should test the old model before replacing it. Method note: the tier figures are the producer's own; the estimate for the last ten languages is a wujec.ai calculation from the published averages and assumes each language is weighted equally.