MDL-5008EST.2026 · IDX.975
AudioIn production

Gemini 3.1 Flash TTS

Google DeepMind · USA · 2026

Speech you direct like an actor — stage directions written straight into the prompt.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Gemini 3.1 Flash TTS is Google's current speech-generation model and the part of the Gemini family built for one-way narration rather than conversation. Its distinguishing idea is steerability: instead of picking a voice from a fixed list and accepting whatever reading it gives, the developer writes the delivery into the prompt — slower here, whispered there — using a set of expressive audio tags. That makes it a tool for audiobooks, dubbing and generated podcasts rather than for call handling, which is the job of the Live models. Pricing follows the shape of the rest of the line: one dollar per million tokens of text going in, twenty per million of audio coming out, so the cost sits almost entirely on the audio side. Google offers it as a public preview, free on the free tier. Two older siblings remain in place for teams that need settled behaviour: 2.5 Flash TTS for cheap low-latency work and 2.5 Pro TTS for high-fidelity, structured productions.

#text to speech#voice synthesis#streaming#expressive audio#api only
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review