Gemini 3.1 Flash Live
Google DeepMind · USA · 2026
Google's answer to live voice — audio in, audio out, no transcription in between.
Gemini 3.1 Flash Live is Google's current low-latency conversational model, served over the Live API as a bidirectional audio stream. The design point it shares with OpenAI's realtime line is native audio: the model hears sound and produces sound directly, instead of passing through a speech-to-text step that throws away tone, pace and interruption. Alongside voice it also accepts images and video, which makes it the model behind camera-based assistants that comment on what they are shown. Where it differs sharply from its American rival is price — three dollars per million audio input tokens and twelve per million output, roughly half a cent and under two cents per minute respectively, an order of magnitude below GPT-Realtime-2. It is offered as a public preview, free of charge on Google's free tier, which in practice means anyone can use it but Google reserves the right to change it. Its production-grade predecessor, the 2.5 Flash native-audio model from December 2025, remains available for workloads that need stability.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!