MDL-2920EST.2025 · IDX.408
Language modelIn production

Gemini 2.0 Flash

Google DeepMind · United States · 2025

Google's cheap workhorse that could see and speak in real time — and the model that made streaming audio and video an ordinary API call.

wujec.ai score

8.4/10

Community score

no votes yet
Sign in to rate

Gemini 2.0 Flash was announced on 11 December 2024 as an experimental release and became generally available on 5 February 2025. Google presented it as the opening of an agentic era, but its lasting contribution was more concrete: it made real-time multimodality cheap and routine. Through the Multimodal Live API a developer could stream audio and video into the model and receive spoken answers back with low latency, so an application could watch a camera feed and hold a conversation about it. The model also generated images natively and produced multilingual speech, rather than handing those jobs to separate systems. Native tool use — Google Search, code execution and user-defined functions — was built in rather than bolted on. The economics were the other half of the story. At roughly 0.10 dollars per million input tokens and 0.40 per million output, with a one-million-token context window, Flash was cheap enough to sit inside high-volume products, and it became Google's default model across the Gemini app and much of Workspace. Google reported that it outperformed the previous generation's Pro model while running at about twice its speed. The model has since been retired. Google's deprecation schedule set 1 June 2026 as the shutdown date for gemini-2.0-flash and gemini-2.0-flash-001, with gemini-3.6-flash as the recommended replacement, and that date has now passed — requests to the 2.0 Flash identifiers no longer work.

#historic#multimodal#realtime#retired
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review