MiniMax M3
MiniMax · China · 2026
Open-weight mixture-of-experts model built around a new sparse attention scheme that makes a one-million-token context affordable.
MiniMax M3, released on 31 May 2026, is the Shanghai lab's flagship language model and its argument that long context should be treated as a dimension to scale rather than a headline number. Its central piece is MSA (MiniMax Sparse Attention), an attention architecture designed in-house to sidestep the quadratic cost of full attention by partitioning the key-value cache into blocks and gathering only the queries that hit each one. MiniMax reports that at a one-million-token context the per-token compute is a twentieth of its previous generation, with prefill more than nine times faster — the practical claim being not that the window exists, but that filling it is not ruinous. The model is natively multimodal: it accepts images and video alongside text, and is positioned for computer-use and agentic work rather than chat alone. Architecturally it is a mixture of experts totalling roughly 427 billion parameters across 60 layers, routing four of 128 experts per token, with a maximum position count of 1,048,576. MiniMax pairs it with MiniMax Code, an agent harness trained together with the model and built on the open-source OpenCode and Pi projects. Weights are published on Hugging Face under MiniMax's own model licence. On the company's API the model is billed by input length: calls up to 512K input tokens sit at one rate, longer calls at double, with a standing 50 percent discount that puts the short-context tier at $0.30 per million input tokens and $1.20 per million output. Subscription tiers from $20 to $120 a month bundle multi-billion-token monthly quotas shared across MiniMax's text, image, speech and music models.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!