Gemini 3.1 Flash TTS
Google DeepMind · USA · 2026
Speech you direct like an actor — stage directions written straight into the prompt.
Gemini 3.1 Flash TTS is Google's current speech-generation model and the part of the Gemini family built for one-way narration rather than conversation. Its distinguishing idea is steerability: instead of picking a voice from a fixed list and accepting whatever reading it gives, the developer writes the delivery into the prompt — slower here, whispered there — using a set of expressive audio tags. That makes it a tool for audiobooks, dubbing and generated podcasts rather than for call handling, which is the job of the Live models. Pricing follows the shape of the rest of the line: one dollar per million tokens of text going in, twenty per million of audio coming out, so the cost sits almost entirely on the audio side. Google offers it as a public preview, free on the free tier. Two older siblings remain in place for teams that need settled behaviour: 2.5 Flash TTS for cheap low-latency work and 2.5 Pro TTS for high-fidelity, structured productions.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!