Qwen3-ForcedAligner-0.6B
Alibaba Cloud · China · 2026
A narrow tool that does one job: matching an existing transcript to the sound, word by word, two to four times more accurately than the toolkit subtitlers use today.
Qwen3-ForcedAligner-0.6B is the least glamorous and possibly the most practical model in Alibaba's January 2026 speech release. It does not transcribe and it does not speak. Given a recording and a text that is already known to match it, it says at what millisecond each word, character or phrase begins and ends. That is the operation behind subtitle timing, karaoke, audiobook-to-ebook syncing, dubbing and the preparation of training data for other speech models — work that is normally done with a patchwork of older alignment tools. Against those tools the numbers are not close. Measured as average absolute deviation of the predicted boundary from a hand-labelled one, it lands 37.5 ms off in English where WhisperX — the alignment layer most subtitling pipelines reach for — is 92.1 ms off, 41.7 against 145.3 in French, 46.5 against 165.1 in German, 33.1 in Chinese where the older Monotonic-Aligner is 161.1. It also covers Japanese, Korean and Portuguese, where the comparison tools produce no result at all. The limits are honest and narrow: eleven languages (Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish), speech only — no singing — and segments of up to five minutes at a time, so longer material has to be cut into chunks. The model is non-autoregressive, which is why it is fast, and at 0.92 billion parameters in the weight files it runs on modest hardware. Apache 2.0, and it plugs straight into Qwen3-ASR as an optional timestamp stage.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!