Qwen3-ASR
Alibaba Cloud · China · 2026
Alibaba's open-weight transcriber: 30 languages and 22 Chinese dialects, streaming and offline in one model, and it beats Whisper everywhere except the rarest tongues.
Qwen3-ASR is the transcription half of Alibaba's January 2026 speech release: two open-weight models, 1.7B and 0.6B, both built on the audio understanding of the Qwen3-Omni foundation model and published under Apache 2.0. They identify the spoken language and transcribe it across 30 languages — Polish among them — plus 22 Chinese dialects and English spoken with accents from several countries. Unusually for this class, one and the same model serves both offline transcription of long recordings and live streaming, and the producer trains it not only on speech but on singing and on songs with backing music, which most transcribers refuse outright. On the standard public benchmarks the 1.7B version edges past OpenAI's Whisper-large-v3, which has been the open default since 2023: 4.90 against 5.27 word error rate on Fleurs, 9.18 against 10.77 on CommonVoice, 8.55 against 8.62 on MLS. Its language identification is a clearer win, 97.9 per cent average accuracy against 94.1. The advantage narrows the further one goes from the major languages, and on the ten rarest in the test set — Czech, Danish, Greek, Persian, Finnish, Filipino, Hungarian, Macedonian, Romanian and Swedish — Whisper still wins by a wide margin, 8.16 against 12.60. The smaller 0.6B model trades accuracy for throughput and is meant for serving many streams at once. Switching either model from offline to streaming costs accuracy too: the 1.7B average rises from 2.69 to 3.33. One caveat on the naming — the published weight files hold 2.35 billion parameters for the model called 1.7B and 0.94 billion for the one called 0.6B, because the audio encoder is counted separately from the name.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!