Voxtral TTS, published on 23 March 2026, is Mistral's first speech-generation model and the mirror image of the Voxtral transcription family: instead of listening, it speaks. Four billion parameters are split across a language decoder, an acoustic transformer and a neural audio codec, and the model streams speech as it is being generated rather than waiting for the sentence to end — which is what makes it usable in live conversation. Given five to twenty-five seconds of reference audio it reproduces a speaker's accent and intonation without any fine-tuning. The weights are public, but unlike the rest of the Voxtral line they carry a non-commercial licence: businesses need a separate agreement with Mistral or must use the paid API.
#text to speech#voice cloning#open weights#streaming#multilingual