MDL-9099EST.2026 · IDX.665
VideoIn production

MiniMax H3

MiniMax · China · 2026

Open-weight omni-modal model that turns mixed text, image, video and audio prompts into 2K clips with native stereo sound.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

MiniMax H3, launched on 31 July 2026 and known on the consumer Hailuo platform as Hailuo 3.0, is the Shanghai company's third-generation video model and the first it released with open weights. Its defining idea is the removal of task boundaries: where earlier systems split generation into separate experts for text-to-video, first-and-last-frame, subject reference, motion transfer and editing, H3 reads text, images, video and audio as a single context and works out from plain language how those materials relate to the clip it is asked to produce. Output runs 4 to 15 seconds at 24 fps with 32 kHz stereo audio — dialogue, music, effects and ambience generated jointly rather than dubbed on afterwards — in aspect ratios from 21:9 to 9:16, with stable spoken dialogue in eleven languages. The system ships in three parts. H3-Context-IR interprets the user's mixed inputs and compiles them into a structured intermediate representation; H3-Base generates video and sound at a 768-pixel short side; H3-Regenerate-2K then feeds that draft back through the base model together with the original context to rebuild it at 2K, which MiniMax says recovers fine detail and small text that a conventional upscaler can only guess at. H3-Base is a 33-billion-parameter dense transformer, deliberately simpler than the Hailuo 02 architecture it replaces, paired with a frozen Qwen3-VL-32B encoder. Weights for the generative modules were published on 3 August 2026 under MiniMax's own Community License, with support in vLLM, diffusers and ComfyUI. The orchestration layer, H3-Context-IR, was not open-sourced: it depends on hosted models, so self-hosted users either call MiniMax's API for it or build their own context pre-processor from the published prompting guide. On the company's own platform 2K video costs $0.13 per second and 768p $0.08 per second, with generated audio billed at nothing.

#video generation#open weights#native audio#omni-modal
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review