MiniMax Speech 2.8
MiniMax · China · 2026
Speech synthesis whose selling point is imperfection: it breathes, hesitates and clears its throat on cue.
MiniMax Speech 2.8, released on 11 August 2026, is the Shanghai company's text-to-speech model and an unusually direct admission of where synthetic voice still gives itself away. The company's own framing is that earlier AI voices sounded cold because they were too perfect; 2.8's headline feature, native sound tags, models the fillers real speech is full of — "um", "uh", breaths, chuckles, a cleared throat — as first-class elements of the text rather than post-processing effects, so the surrounding rhythm, pitch and pauses shift around them. The second claim concerns cloning. MiniMax says a ten-second sample is enough to capture a speaker's texture, breathiness and pace, and the pricing reflects an assumption of casual use: a cloned voice costs $1.50, a designed one $3. Synthesis itself is billed per character, $60 per million for the turbo tier and $100 per million for HD. The third strand is engineering rather than expression — a reworked processing chain aimed at removing background noise and synthetic artefacts from the output. On multilinguality the company is deliberately narrow. Rather than announcing a language count, it says it has fixed "accent bleed" for one pair, Mandarin and Japanese, where tones and pronunciation used to drift, and promises the same treatment for further languages without naming them or giving dates. The model is closed: there are no weights, only the MiniMax Audio product and the platform API.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!