Voxtral Small is Mistral AI's flagship audio model, released in July 2025 as the company's first move into speech. Unlike a pure transcription engine, it is built on the Mistral Small 3.1 language backbone, so a single call can turn thirty minutes of recording into text — or answer questions about it, summarise it and trigger functions straight from the spoken instruction, without a separate language model in the loop. Mistral publishes it under the Apache 2.0 licence, which makes it one of the few frontier-grade audio models a company can run entirely on its own hardware. On Mistral's own benchmarks it outperforms Whisper large-v3 across transcription tasks and holds up against the paid transcription tiers of the large American labs.