MDL-3542EST.2026 · IDX.067
Language modelIn production

Ming-flash-omni 2.0

InclusionAI (Ant Group) · China · 2026

An open-weight model that takes text, images, video and sound in, and gives text, images and sound back — including cloned voices — on 6B active parameters.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Ming-flash-omni 2.0 is the multimodal line of InclusionAI, the open-source arm of Ant Group, and the widest-ranging model the publisher has released: it accepts image, text, video and audio, and produces image, text and audio. Ant announced the finished version on 11 February 2026, a day after the weights appeared on Hugging Face. It sits on the same Ling 2.0 mixture-of-experts backbone as the publisher's text models — 100 billion total parameters with about 6 billion active per token, 104.23 billion counted in the published safetensors index. That activation ratio is the commercial argument: an omni model of this reach that only lights up 6 billion parameters per token is far cheaper to serve than a dense equivalent. The sound side is the most unusual part. Speech, general audio and music are generated through one pipeline rather than three, using continuous autoregression with a diffusion transformer head, and the model does zero-shot voice cloning with separate control over emotion, timbre and background atmosphere. On the image side, segmentation, generation and editing share a single architecture, which the maker aims at tasks such as removing an object while keeping texture and depth consistent. It also handles streaming video conversation — sustained dialogue over a live camera feed rather than single-shot analysis. The line has a long history for such a young field: a test build in May 2025, Ming-lite-omni v1 later that month, v1.5 in July, a Ming-flash-omni preview in October 2025, and this release. Two technical reports are published, from June and October 2025. As with the Ling-1T flagship, a reader should know that the model card contains no benchmark table at all — the claim of state-of-the-art results among open-source omni models is made in prose and illustrated with images, with no figures to check. The licence label in the metadata reads MIT, and once again the repository ships no licence file among its 98 files, and the card text has no licence section. Downloads run at about 2,300 in the 30 days to 29 August 2026.

#open weights#omni#MoE#MIT#image generation#voice cloning
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review