All newsResearch

The best model at telling speakers apart still gets three words in ten wrong — and the standard metric hid it

Published: 8/29/2026 · Source: AssemblyAI

AssemblyAI has published the launch material for Universal-3.5 Pro, its flagship speech-to-text model, and the most interesting part is not a capability claim. It is an argument that the measure the industry has used for years to judge who-said-what is the wrong measure. The field has long reported diarisation error rate, DER. DER compares time regions: it asks which stretches of the recording were attributed to which speaker. It never looks at the words. AssemblyAI's objection is that this can rank systems backwards. A transcript in which every single word is on the right speaker can still score above 30% DER — because the system declined to label laughter as speech, or because the hand-annotated speech segments are simply looser than the word timestamps a modern model returns. The company has switched to cpWER, concatenated minimum-permutation word error rate. That measure asks the question a reader would ask: of everything this speaker said, what fraction did the system get wrong or credit to somebody else? Measured that way, the numbers are far less flattering than a decade of DER charts would suggest. AssemblyAI's own new model posts the best result in its comparison — an average cpWER of 30.17 across six benchmark collections. That is roughly three words in ten wrong or misattributed. The spread matters: on telephone audio the model records 17.18 and 17.78, on meeting recordings 27.36, but on the harder conversational sets 37.02 and 48.22. Every rival in the table scores worse on average, from 35.26 to 44.58, though these are the manufacturer's measurements of its competitors and should be read as such. The timing sharpens the point. Two days ago Google released Gemini 3.5 Transcribe, whose live endpoint does not do speaker attribution at all, and whose file endpoint halves the maximum recording length when diarisation is switched on and calls attribution above three speakers experimental. Read next to AssemblyAI's table, that looks less like a product gap and more like an honest reflection of how unsolved the problem is. There is a caveat worth stating plainly: a vendor changing the yardstick it is judged by is always worth a second look, and AssemblyAI's new model happens to top the new yardstick. But the criticism of DER stands on its own — it is a measure of time, and readers of a transcript care about words. The new number is less flattering to everyone in the industry, including the company publishing it. Universal-3.5 Pro was released on 29 June 2026 and becomes AssemblyAI's default model for recorded audio on 2 September 2026, when the previous Universal-3 Pro line is retired. It costs 0.21 dollars per hour of audio; the streaming variant costs 0.45 dollars per hour of open session. The model now has a profile in the catalogue, as does its manufacturer, which wujec.ai had not covered until today.