Mercury 2.5
Inception · United States · 2026
A language model that does not write word by word: Inception's diffusion LLM refines whole blocks of tokens in parallel and reports 1,107 tokens per second on ordinary NVIDIA cards.
Mercury 2.5 is the current flagship of Inception, an American company built around a bet that almost nobody else in the industry has taken into production: that a large language model does not have to generate text one token after another. Mercury models are diffusion language models — dLLMs — which start from a rough draft of a whole block of output and refine it over several passes, producing many tokens in parallel. Inception says Mercury 2.5 is, to its knowledge, the largest diffusion language model ever trained. The company was founded by researchers from Stanford, UCLA and Cornell, with engineers from Google DeepMind, Meta AI, Microsoft AI and OpenAI. The point of the architecture is time, not intelligence. Inception publishes 1,107 tokens per second on widely available NVIDIA GPUs, and the whole product line is aimed at work where a pause is a defect rather than an inconvenience: voice agents that hold a live telephone call, search pipelines that fire dozens of model calls inside one user request, and coding assistants that constantly compact and re-route their own context. The customer numbers the company cites come from that world — OpenCall reports a median model response near 170 milliseconds on live calls, and Augment Code reports context compaction dropping from roughly 150 seconds to 27 with a 90 per cent cost reduction. All of those figures are reported by the vendor and its customers, not measured independently. On quality Inception is careful not to claim a frontier position. It describes Mercury 2.5 as 40 per cent more intelligent than Mercury 2 and comparable to the cost-optimised tier of the big labs — GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, Claude Haiku 4.5 — which places it as a fast, cheap workhorse rather than a reasoning champion. The model takes 260K tokens of context, emits up to 65,536, and supports tunable reasoning effort, parallel tool calls and schema-aligned JSON through an OpenAI-compatible API. The price needs a footnote. Inception's list price is $0.20 per million input tokens and $0.75 per million output, with cached input at $0.02. At launch the company put an 80 per cent discount on the model, so through September 2026 it bills $0.04, $0.15 and $0.004 respectively, and those are the promotional numbers that appear on third-party listings. Weights are not published; the model runs through the Inception API, Baseten and OpenRouter, with dedicated capacity for enterprise deployments. Two siblings were announced as previews on the same day: Mercury Voice, a 128K dLLM tuned for voice agents with time to first token under 170 milliseconds, and Mercury Router, which reads an incoming prompt and dispatches it to whichever open or closed model fits best. Neither has published pricing. Older members of the line — Mercury 1, Mercury 2 and Mercury Edit 2 — are, in the maker's own words, supported for existing customers only, which in practice closes them to new users while leaving them running.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!