Voxtral Mini Transcribe 2
Mistral AI · Francja · 2026
Mistral's batch transcription model — speaker diarization, word-level timestamps and 3-hour recordings in 13 languages, at $0.003 per minute.
Voxtral Mini Transcribe 2, released on 4 February 2026, is the batch half of Mistral's second-generation speech-to-text line. It is built for material that already exists — meeting recordings, call-centre archives, interviews — rather than for live conversation, and it accepts up to three hours of audio in a single request. What separates it from a plain transcription endpoint is what it returns alongside the words. It performs speaker diarization, marking who spoke when, which is the difference between an unusable wall of text and a readable transcript of a four-person meeting. It emits timestamps at the level of individual words, so a transcript can be aligned back to the recording. And it supports context biasing with up to 100 custom terms, which is how product names, drug names or surnames stop being transcribed phonetically. Coverage is 13 languages: English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian and Dutch. Mistral prices it at $0.003 per minute of audio. Its sibling, Voxtral Mini Transcribe Realtime, handles live streams instead and has its own profile; diarization is not available in the realtime model. The first-generation Voxtral Mini Transcribe endpoint was retired on 31 May 2026.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!