MDL-6647EST.2026 · IDX.103
AudioIn production

Gemini 3.5 Transcribe

Google DeepMind · USA / UK · 2026

Google's dedicated speech-to-text model, sold as two endpoints that do not do the same things: the file version transcribes an hour and names the speakers, the live version does neither.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Gemini 3.5 Transcribe is Google's first speech-to-text model sold on its own rather than as a side capability of a general Gemini model. It reached general availability on 26 August 2026 in two versions that share a name and a price list but not a feature set. The file version, gemini-3.5-transcribe, takes recordings of up to an hour, detects the language by utterance across more than 85 languages including Polish, follows speakers who switch languages mid-sentence, separates up to eight speakers and marks timestamps on individual words. Two of those features come at a cost stated by Google itself: with diarization or word timestamps switched on, the audio limit falls from an hour to thirty minutes, and Google notes that word-level timestamps degrade transcription accuracy. Attribution beyond three speakers is described as experimental. The live version, gemini-3.5-transcribe-live, streams over WebSockets, returning interim results before final ones and supporting several voice-activity-detection strategies. It cannot name speakers and cannot place word timestamps, and a session lasts ten minutes. Custom vocabulary biasing of up to 1,000 terms works on both, though Google says results are usually best with no more than a hundred. The price runs against intuition. The restricted live endpoint is the more expensive one: about $0.009 per minute of audio on Google's own blended estimate, against about $0.005 for the file endpoint that does more. In list terms the file version costs $2.00 per million audio input tokens and $12.00 per million text output tokens, the live version $3.50 and $21.00. Both have a free tier, on which Google states that requests are used to improve its products; the paid tier states they are not. Neither version supports caching, code execution, file search, function calling or the thinking mode that Gemini's general models use. This is a transcription tool, not a conversational one.

#speech-to-text#diarization#streaming#85+ languages#low latency
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review