MDL-7680EST.2026 · IDX.902
Language modelIn production

Mercury 2

Inception · United States · 2026

The diffusion model that made Inception's architecture a reasoning engine: Mercury 2 reports 1,009 tokens per second on NVIDIA Blackwell cards and is now sold only to customers who already had it.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Mercury 2 is the model with which Inception moved its diffusion approach out of the coding niche and into general reasoning. Earlier Mercury releases were pitched at autocomplete and chat; this one was announced as the world's fastest reasoning language model, and the argument behind it was about the shape of production workloads rather than raw benchmark scores. Inception's case is that production AI is no longer one prompt and one answer but loops — agents, retrieval pipelines, extraction jobs — in which latency does not appear once but compounds across every step, every user and every retry. The architecture is the same bet the whole company rests on. Instead of decoding one token after another from left to right, the model refines a full draft across all positions at once and converges over a small number of passes. Inception describes it as less typewriter, more editor revising a whole page. The published number is 1,009 tokens per second on NVIDIA Blackwell GPUs, with a claimed speed advantage above five times over sequential decoding; NVIDIA supplied a supporting quote, which is worth reading as what it is — a vendor endorsement, not an independent measurement. The practical reading of Mercury 2 is a trade-off rather than a victory. Inception does not claim frontier quality for it, only that it is competitive with leading speed-optimised models, and the interesting claim is a different one: that diffusion-based reasoning delivers reasoning-grade answers inside a real-time latency budget, where an autoregressive model would have to buy quality with longer chains and more retries. The model takes 128K tokens of context, emits up to 50,000, and offers tunable reasoning effort, native tool use and schema-aligned JSON through an OpenAI-compatible API. Pricing is $0.25 per million input tokens, $0.75 output and $0.025 for cached input. Weights are not published. Its status now needs care, because the company's two pages disagree in tone. The models page carries a footnote saying that Mercury 1, Mercury 2 and Mercury Edit 2 remain supported for existing customers, with access and migration handled by a sales representative — in practice a closed door for new accounts. The developer documentation, meanwhile, still lists Mercury 2 with a full price table, a 128K context window and the chat completions endpoint. So the model is running and billable, but the route in is through an account that already exists, and the successor announced in September 2026, Mercury 2.5, is what the company sells today at a quarter of the input price and twice the context.

#proprietary#diffusion LLM#low latency#reasoning#tool use#128K context#USA
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review