Mercury Voice
Inception · United States · 2026
A diffusion model sold on one number: under 170 milliseconds to the first token. It is the pause a telephone caller hears, and Inception is not yet quoting a price for removing it.
Mercury Voice is the branch of Inception's diffusion line built for a single constraint: a voice agent cannot think in silence. Announced as a preview on 8 September 2026 alongside the Mercury 2.5 flagship, it is described by the maker as a dLLM optimised for voice agents with the tightest latency budgets, and it is sold on one figure — time to first token under 170 milliseconds. That number is worth unpacking, because it is not the same claim as tokens per second. In a live telephone call, throughput matters less than the gap before the machine starts talking: past roughly 300 milliseconds a caller begins to hear a pause and treat it as hesitation. Inception's argument, stated in its launch post, is that in voice latency is not an infrastructure detail but the pause the caller hears. The company supports it with a customer figure rather than a lab measurement — OpenCall, which runs AI phone agents on live customer calls, reports median model response latency close to 170 milliseconds on its production workload. It is a vendor-reported number from a vendor's customer. What the model is good at beyond speed is not documented. Inception's models page gives it a 128K context window and the same feature set as the flagship — reasoning, tool use, structured output — and names customer support, patient care, education and gaming as target uses. There is no benchmark table, no parameter count, no published quality comparison against the general Mercury models, and no weights. The commercial status is the honest headline. Unlike Mercury 2.5, which has a public price list, Mercury Voice has no published rate at all: the pricing section of the maker's page is a link to contact sales, and Inception offers instead to work with prospective users on testing workload fit under their own serving constraints. So this is a real product being fitted to customers one at a time, not a model an account holder can call this afternoon — which is why this catalogue lists it as a pilot rather than in production. Anyone comparing it with an off-the-shelf voice stack should treat both the latency figure and the absence of a price as part of the same sentence.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!