All newsReleases

OpenAI's smartest model now answers 14 times faster — on chips OpenAI did not build

Published: 8/15/2026 · Source: OpenAI, Previewing Ultrafast mode

OpenAI announced on 13 August 2026 an early preview of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second — up to 14 times faster than standard processing. The company has published no price and no general availability date. Access is limited to a selected group of customers, with a sign-up form for everyone else. The engineering is not OpenAI's. Ultrafast runs on Cerebras wafer-scale processors, which hold 44 GB of SRAM on the chip itself and therefore sidestep the memory-bandwidth ceiling that limits token generation on conventional GPU clusters. That is the fact worth pausing on: the most capable model OpenAI sells reaches its highest speed on silicon designed by another company, and neither NVIDIA nor OpenAI's own hardware programme is in the sentence. OpenAI frames the change as a shift in what has to be traded away. Until now, real-time response meant reaching for a smaller or more specialised model; the company's phrase for the alternative is "more useful work per second". The uses it names are ones where the clock is part of the problem — reading logs and recent commits while an outage is still unfolding, assessing transactions while market conditions move, answering a shopper before hesitation turns into an abandoned cart. Internally, OpenAI says research batches that used to run overnight now fit inside a working day as several iterations. One widely repeated comparison should be read with care. The figures showing GPT-5.6 Sol on Ultrafast completing all 2,500 questions of Humanity's Last Exam in 11 hours 11 minutes against 78 hours 27 minutes for Claude Fable 5 come from Cerebras, the hardware supplier, not from an independent evaluator — and they measure wall-clock time for a benchmark run, not accuracy. OpenAI's own claim is narrower and more useful: on GDP-Val, its benchmark of economically valuable knowledge work, it reports a 5.6-fold end-to-end speedup over standard processing with no loss of quality. Ultrafast is a service tier, not a new model. The weights are the same GPT-5.6 Sol already in the catalogue; what changed is where they run and how fast the tokens come out.