MDL-1346EST.2026 · IDX.112
AudioIn production

MiniMax Speech 2.8

MiniMax · China · 2026

Speech synthesis whose selling point is imperfection: it breathes, hesitates and clears its throat on cue.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

MiniMax Speech 2.8, released on 11 August 2026, is the Shanghai company's text-to-speech model and an unusually direct admission of where synthetic voice still gives itself away. The company's own framing is that earlier AI voices sounded cold because they were too perfect; 2.8's headline feature, native sound tags, models the fillers real speech is full of — "um", "uh", breaths, chuckles, a cleared throat — as first-class elements of the text rather than post-processing effects, so the surrounding rhythm, pitch and pauses shift around them. The second claim concerns cloning. MiniMax says a ten-second sample is enough to capture a speaker's texture, breathiness and pace, and the pricing reflects an assumption of casual use: a cloned voice costs $1.50, a designed one $3. Synthesis itself is billed per character, $60 per million for the turbo tier and $100 per million for HD. The third strand is engineering rather than expression — a reworked processing chain aimed at removing background noise and synthetic artefacts from the output. On multilinguality the company is deliberately narrow. Rather than announcing a language count, it says it has fixed "accent bleed" for one pair, Mandarin and Japanese, where tones and pronunciation used to drift, and promises the same treatment for further languages without naming them or giving dates. The model is closed: there are no weights, only the MiniMax Audio product and the platform API.

#text to speech#voice cloning#sound tags#closed API
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review