MDL-5437EST.2026 · IDX.711
category.language-modelIn production

Gemini 3.5 Flash-Lite

Google · USA · 2026

Google's cheapest current Gemini tier — multimodal input at $0.30 per million tokens, generally available since 21 July 2026.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Gemini 3.5 Flash-Lite is the bottom price tier of Google's 3.5 generation, promoted to general availability on 21 July 2026 together with Gemini 3.6 Flash. Google positions it as its fastest and most cost-effective 3.5 model for high-throughput execution, and in practice that means two jobs: bulk work where volume matters more than depth — classification, extraction, routing, summarising queues of documents — and the subagent role, where a larger model plans and a swarm of cheap calls does the legwork. The economics are the point. Paid usage costs $0.30 per million input tokens, covering text, image, video and audio alike, and $2.50 per million output tokens; context caching is $0.03 per million tokens plus $1.00 per million tokens per hour of storage. That is five times cheaper on input than Gemini 3.5 Flash ($1.50 / $9.00) for the same kinds of input. There is a free tier for both input and output, which makes the model the natural first stop for anyone testing an idea before paying for it. It is multimodal on the way in — text, images, video and audio — and text on the way out. Google has not published a separate specification table for the Lite variants on its public model overview, so this profile does not state token limits or a knowledge cutoff; we will fill them in when the model card appears. Its predecessor, Gemini 3.1 Flash-Lite, remains available and has a published shutdown date of 7 May 2027, with 3.5 Flash-Lite named as the replacement.

#multimodal#low-cost#high-throughput#subagent
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review