News

What's happening in robotics and AI — curated by the wujec.ai editors.

Business8/18/2026 · OpenAI — cennik API i komunikat o trybie Ultrafast (obliczenia własne wujec.ai)

OpenAI's price ladder for its flagship doubles at every rung — and the only tier with a published speed is the one you cannot buy

Five days after OpenAI published the first tokens-per-second figure in the history of its API — 750 output tokens per second for the Ultrafast preview of GPT-5.6 Sol — that tier still has no price. The tier that does have a price still has no speed. The two facts belong together, and the shape of OpenAI's price list makes the point sharper than either announcement does. OpenAI sells GPT-5.6 Sol at three priced service levels, and on its own pricing page each level is exactly double the one below it. Flex and Batch processing cost USD 2.50 per million input tokens and USD 15 per million output. Standard costs USD 5 and USD 30. Fast mode costs USD 10 and USD 60. The doubling also holds on the long-context rows, which apply to requests above the model's 272,000-token threshold: 5 and 22.50 for Flex, 10 and 45 for Standard, 20 and 90 for Fast. Eight numbers, four exact factors of two, no rounding and no exceptions. What the ladder does not carry is a single speed. Fast mode carries a 100% premium over Standard, and OpenAI has never published a tokens-per-second rate, a latency figure or a percentage to say what that premium delivers. Flex is documented only as offering "slower response times and occasional resource unavailability" — again with no number attached. A customer choosing between the rungs is choosing between prices that are precise to the cent and speeds that are not stated at all. Ultrafast inverts this exactly. It is the first service level OpenAI has attached a throughput figure to — "up to 750 output tokens per second", "up to 14× faster than Standard processing" — and it is the only one with no price, available in limited preview to a select group of customers. One baseline can be derived from those two figures, with a caveat that matters. If the 750 tokens per second and the 14× multiplier describe the same run, Standard processing generates roughly 54 output tokens per second. That would be the closest thing to a published baseline this model has. But both figures are ceilings marked "up to", and OpenAI does not state that they were measured on the same run, so 54 is an inference from the announcement rather than a number OpenAI has disclosed. What this does not prove: nothing here indicates what Ultrafast will cost when it is priced. The doubling across the three existing rungs describes the current list, not a commitment about the next one, and a preview price need not survive to general availability.

GPT-5.6 Sol →
Releases8/15/2026 · OpenAI, Previewing Ultrafast mode

OpenAI's smartest model now answers 14 times faster — on chips OpenAI did not build

OpenAI announced on 13 August 2026 an early preview of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second — up to 14 times faster than standard processing. The company has published no price and no general availability date. Access is limited to a selected group of customers, with a sign-up form for everyone else. The engineering is not OpenAI's. Ultrafast runs on Cerebras wafer-scale processors, which hold 44 GB of SRAM on the chip itself and therefore sidestep the memory-bandwidth ceiling that limits token generation on conventional GPU clusters. That is the fact worth pausing on: the most capable model OpenAI sells reaches its highest speed on silicon designed by another company, and neither NVIDIA nor OpenAI's own hardware programme is in the sentence. OpenAI frames the change as a shift in what has to be traded away. Until now, real-time response meant reaching for a smaller or more specialised model; the company's phrase for the alternative is "more useful work per second". The uses it names are ones where the clock is part of the problem — reading logs and recent commits while an outage is still unfolding, assessing transactions while market conditions move, answering a shopper before hesitation turns into an abandoned cart. Internally, OpenAI says research batches that used to run overnight now fit inside a working day as several iterations. One widely repeated comparison should be read with care. The figures showing GPT-5.6 Sol on Ultrafast completing all 2,500 questions of Humanity's Last Exam in 11 hours 11 minutes against 78 hours 27 minutes for Claude Fable 5 come from Cerebras, the hardware supplier, not from an independent evaluator — and they measure wall-clock time for a benchmark run, not accuracy. OpenAI's own claim is narrower and more useful: on GDP-Val, its benchmark of economically valuable knowledge work, it reports a 5.6-fold end-to-end speedup over standard processing with no loss of quality. Ultrafast is a service tier, not a new model. The weights are the same GPT-5.6 Sol already in the catalogue; what changed is where they run and how fast the tokens come out.

GPT-5.6 Sol →