MDL-9919EST.2026 · IDX.612
AudioIn production

Gemini 3.1 Flash Live

Google DeepMind · USA · 2026

Google's answer to live voice — audio in, audio out, no transcription in between.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Gemini 3.1 Flash Live is Google's current low-latency conversational model, served over the Live API as a bidirectional audio stream. The design point it shares with OpenAI's realtime line is native audio: the model hears sound and produces sound directly, instead of passing through a speech-to-text step that throws away tone, pace and interruption. Alongside voice it also accepts images and video, which makes it the model behind camera-based assistants that comment on what they are shown. Where it differs sharply from its American rival is price — three dollars per million audio input tokens and twelve per million output, roughly half a cent and under two cents per minute respectively, an order of magnitude below GPT-Realtime-2. It is offered as a public preview, free of charge on Google's free tier, which in practice means anyone can use it but Google reserves the right to change it. Its production-grade predecessor, the 2.5 Flash native-audio model from December 2025, remains available for workloads that need stability.

#speech to speech#realtime#voice agents#native audio#streaming
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review