Mercury 2
Inception · United States · 2026
The diffusion model that made Inception's architecture a reasoning engine: Mercury 2 reports 1,009 tokens per second on NVIDIA Blackwell cards and is now sold only to customers who already had it.
Mercury 2 is the model with which Inception moved its diffusion approach out of the coding niche and into general reasoning. Earlier Mercury releases were pitched at autocomplete and chat; this one was announced as the world's fastest reasoning language model, and the argument behind it was about the shape of production workloads rather than raw benchmark scores. Inception's case is that production AI is no longer one prompt and one answer but loops — agents, retrieval pipelines, extraction jobs — in which latency does not appear once but compounds across every step, every user and every retry. The architecture is the same bet the whole company rests on. Instead of decoding one token after another from left to right, the model refines a full draft across all positions at once and converges over a small number of passes. Inception describes it as less typewriter, more editor revising a whole page. The published number is 1,009 tokens per second on NVIDIA Blackwell GPUs, with a claimed speed advantage above five times over sequential decoding; NVIDIA supplied a supporting quote, which is worth reading as what it is — a vendor endorsement, not an independent measurement. The practical reading of Mercury 2 is a trade-off rather than a victory. Inception does not claim frontier quality for it, only that it is competitive with leading speed-optimised models, and the interesting claim is a different one: that diffusion-based reasoning delivers reasoning-grade answers inside a real-time latency budget, where an autoregressive model would have to buy quality with longer chains and more retries. The model takes 128K tokens of context, emits up to 50,000, and offers tunable reasoning effort, native tool use and schema-aligned JSON through an OpenAI-compatible API. Pricing is $0.25 per million input tokens, $0.75 output and $0.025 for cached input. Weights are not published. Its status now needs care, because the company's two pages disagree in tone. The models page carries a footnote saying that Mercury 1, Mercury 2 and Mercury Edit 2 remain supported for existing customers, with access and migration handled by a sales representative — in practice a closed door for new accounts. The developer documentation, meanwhile, still lists Mercury 2 with a full price table, a 128K context window and the chat completions endpoint. So the model is running and billable, but the route in is through an account that already exists, and the successor announced in September 2026, Mercury 2.5, is what the company sells today at a quarter of the input price and twice the context.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!