MDL-4363EST.2026 · IDX.969
AudioIn production

Qwen3-TTS

Alibaba Cloud (Qwen / Tongyi Lab) · China · 2026

Alibaba's open-weight voice engine: ten languages, a 97 ms first packet and a clone from three seconds of tape.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Qwen3-TTS, open-sourced on 22 January 2026, is the Qwen team's answer to the closed voice engines that dominate the market. Two sizes were published at once — 0.6B for machines with modest graphics memory and 1.7B for the best quality — and both use a multi-codebook design in which speech is emitted as a stream rather than rendered after the sentence ends. The first audio packet arrives in as little as 97 milliseconds, which is the threshold below which a synthetic voice stops sounding like a recording and starts sounding like a reply. Three abilities are bundled together: cloning a speaker from roughly three seconds of reference audio, designing an entirely new voice from a written description of it, and reading with a curated set of premium timbres — nine of them, spread across genders, ages and languages, plus Chinese dialect profiles including Beijing and Sichuan. The weights carry an Apache 2.0 licence, so unlike most rivals in this class the model may be used commercially without a separate agreement, and it runs on a consumer graphics card. Alibaba also sells a hosted variant through Model Studio for those who would rather not run it themselves.

#text to speech#voice cloning#open weights#streaming#multilingual
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review