MDL-3027EST.2025 · IDX.668
Language modelRetired

Mercury Coder

Inception · United States · 2025

The model that took diffusion out of image generation and into working code: two sizes running at 1,109 and 737 tokens per second on ordinary NVIDIA H100s, launched in February 2025 as the first commercial-scale diffusion language model.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Mercury Coder is where Inception's whole catalogue begins. Announced on 26 February 2025 together with the Mercury family, it was presented as the first commercial-scale diffusion large language model — a claim about architecture rather than about benchmark leadership. Every mainstream language model of the time generated text left to right, one token after another, so a long answer was a long wait by construction. Diffusion models, which already powered image, video and audio generation, work the other way round: they start from noise and refine a whole draft over a handful of denoising passes, revising many positions at once. Applying that idea to discrete data such as text and code had failed repeatedly before; Mercury Coder was the first attempt that shipped to paying customers. The payoff was speed measured on ordinary hardware. Inception reported 1,109 tokens per second for Mercury Coder Mini and 737 for Mercury Coder Small on NVIDIA H100 GPUs, against roughly 200 tokens per second for the fastest speed-optimised autoregressive models of the period — throughput that until then had required specialised inference chips from Groq, Cerebras or SambaNova. The company was careful to frame this as an algorithmic gain that would compound, not compete, with faster silicon. Quality was pitched as parity rather than supremacy, and the published table supports that reading. Mercury Coder Mini scored 88.0 on HumanEval, 77.1 on MBPP, 78.6 on EvalPlus, 74.1 on MultiPL-E, 42.0 on BigCodeBench and 82.2 on fill-in-the-middle; the larger Small variant reached 90.0, 76.6, 80.4, 76.2, 45.5 and 84.8 on the same tests. Both scored poorly on LiveCodeBench (17.0 and 25.0), the hardest competitive-programming set in the set — this was a fast assistant for everyday code, not a reasoning champion. The one genuinely striking result is fill-in-the-middle, where the diffusion approach beat every autoregressive model in the comparison by a wide margin: editing inside an existing file is exactly the task a left-to-right decoder is worst suited to. Inception also reported that on Copilot Arena, where developers vote on live completions, Mercury Coder Mini tied for second place while being about four times faster than GPT-4o Mini. The model was sold through an API and on-premise deployments with fine-tuning support, and a free playground hosted with Lambda Labs; list prices were never published, with enterprise access handled by sales. Two follow-ups built directly on it: the Inception API in April 2025 added fill-in-the-middle for autocomplete, and in September 2025 the same model gained apply-edit — merging a generated snippet into an existing file — reported at 92 percent accuracy on Kortix AI's dataset, matching frontier models while running 46 times faster. That capability became a product line of its own with Mercury Edit. Mercury Coder is no longer on sale. Inception's models page today lists only Mercury 2.5, Mercury Voice and Mercury Router, and its legacy footnote names Mercury 1, Mercury 2 and Mercury Edit 2 as still supported for existing customers — Mercury Coder is not named at all. The profile is kept because this is the model that made the diffusion bet in public and every later Mercury is a descendant of it.

#proprietary#diffusion LLM#code generation#low latency#first generation#USA
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review