MDL-8792EST.2025 · IDX.781
AudioIn production

Fun-CosyVoice 3.0

Alibaba (FunAudioLLM) · China · 2025

Alibaba's other open voice engine: nine languages but 18-plus Chinese dialects, and a phoneme-level fix for words it mispronounces.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Fun-CosyVoice 3.0 comes from FunAudioLLM, Alibaba's speech laboratory, and is the third generation of a line that started in 2024 — long before the Qwen team shipped a voice engine of its own. That makes Alibaba a company selling two open text-to-speech models at once, and they are not the same product. Qwen3-TTS, published six weeks later, is broader across world languages and faster off the mark; CosyVoice 3 goes deep into Chinese instead, covering more than eighteen regional dialects and accents — Cantonese, Minnan, Sichuan, Shanghai, Tianjin, Gansu and others — against nine world languages. The feature that distinguishes it in production work is pronunciation inpainting: when the model reads a word wrong, the operator can pin that single word to explicit Chinese pinyin or English CMU phonemes instead of re-recording, rewriting the sentence or retraining anything. Alongside it the model does its own text normalisation — numbers, symbols and formatted strings are read without the separate rule-based front end that older systems require — and it takes plain-language instructions for language, dialect, emotion, speed and volume. Synthesis is bi-streaming: text can arrive in a stream and audio leaves in a stream, with latency down to about 150 milliseconds. Voice cloning is zero-shot, including across languages, so a voice sampled in one language can read another. The weights are 0.5B, Apache 2.0, and a reinforcement-learned variant ships alongside the base model. On Hugging Face it is the most-liked model the laboratory has published.

#text to speech#voice cloning#open weights#streaming#dialects
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review