All newsReleases

MiniMax teaches its voice model to stumble: Speech 2.8 makes breaths and "um" part of the text

Published: 8/11/2026 · Source: MiniMax

The headline feature of MiniMax Speech 2.8 is what the company calls native sound tags. Rather than splicing in a breath or a laugh after synthesis, the model reads markers like "(breath)", "(chuckle)" and "(clear-throat)" inside the script and generates them as part of the utterance, letting the surrounding pace, pitch and pauses adjust around them. Colloquial fillers — "um", "uh", "ah" — are modelled the same way. MiniMax's reasoning, stated plainly in its announcement, is that earlier AI voices felt cold precisely because they were too perfect, and that the imperfections carry the emphasis and emotion. Two other changes accompany it. Voice cloning now claims to capture a speaker's texture, breathiness and speaking pace from a ten-second sample; at $1.50 per cloned voice and $3.00 per designed one, the pricing assumes this is something users do routinely rather than once. Separately, the processing chain has been reworked to strip background noise and synthetic artefacts from the output. On languages the company is unusually restrained. It announces no language count, only that it has corrected "accent bleed" — tones and pronunciation drifting between languages — for a single pair, Mandarin and Japanese, with more promised but neither named nor dated. Synthesis is billed per character, $60 per million in the turbo tier and $100 per million in HD, through a closed API; unlike MiniMax's video model H3, whose weights were published on 3 August, Speech 2.8 ships without weights.