MDL-3202EST.2026 · IDX.348
AudioIn production

Universal-3.5 Pro

AssemblyAI · USA · 2026

A transcription model that decides who is speaking inside the same pass that decides what was said — and that switches language mid-sentence without being told.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Universal-3.5 Pro is AssemblyAI's flagship speech-to-text model, released on 29 June 2026 for pre-recorded audio and, in a streaming variant, for live transcription and the company's Voice Agent API. From 2 September 2026 it becomes the default model for accounts that do not pin a version, and the previous Universal-3 Pro line is retired. The model's defining design choice is that speaker attribution is not a separate system. Conventional pipelines run transcription and diarisation independently and then stitch the two together by matching timestamps, which breaks on short turns, interruptions and crosstalk. Universal-3.5 Pro emits the transcript and the speaker changes from the same pass, so a fast back-and-forth between two people stays a fast back-and-forth rather than one long block reassigned after the fact. The second change is code switching. Eighteen languages are supported at what the manufacturer calls full accuracy, and a sentence may move between them without configuration — an English-Hindi or English-Mandarin conversation is transcribed in the language each word was actually spoken in, rather than collapsed into one. On AssemblyAI's own code-switching benchmark the model records an average word error rate of 7.69% across five language pairs; the same test puts several competing models above 10% and two of them above 35%. These are the manufacturer's measurements of its rivals and should be read as such. The third feature is contextual prompting: the caller can hand the model a paragraph of domain context — a prior clinical note, a meeting agenda, a list of product names — and the model uses it to disambiguate. AssemblyAI reports a 31% drop in missed medical terms in an internal healthcare test when a previous visit note was supplied. Pricing is per hour rather than per token: 0.21 USD per hour of pre-recorded audio, and 0.45 USD per hour of open WebSocket session for streaming, where silence and idle time are billable. The older Universal-2 remains in the catalogue at 0.15 USD per hour and covers 99 languages instead of 18 — a trade the manufacturer states plainly.

#speech-to-text#diarization#streaming#code-switching#18 languages
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review