News
What's happening in robotics and AI — curated by the wujec.ai editors.
Z.ai opened the weights of its cheap model — the flagship it promised two weeks ago is still closed
Z.ai published GLM-5.3-Flash on 26 August 2026 and put the weights on Hugging Face the same week, under a plain MIT licence. Two weeks earlier the company launched its flagship GLM-5.3 with a promise that its weights would follow in about a fortnight. That deadline has now passed, and the flagship repository does not exist: the newest Z.ai model anyone can download is the cheap one. The last flagship whose weights were actually released is GLM-5.2 from June 2026, at 753.3 billion parameters. GLM-5.3-Flash is less than half that size — 320 billion parameters with 18 billion active per token, and the published safetensors index confirms it at 321.3 billion. The smaller model is not merely a trimmed version. Z.ai describes it as the first open-weight frontier model to combine sparse attention with linear attention, and quantifies the gain against GLM-5.3: attention computation down 3.01 times, KV cache down 4.44 times. It is also the first natively multimodal model in the family, meaning it can look at a rendered interface and correct its own code rather than working blind. The price is where that architecture shows. GLM-5.3-Flash lists at 0.15 dollars per million input tokens and 0.50 per million output; GLM-5.3 costs 1.40 and 4.40. That is 9.3 times cheaper on input and 8.8 times cheaper on output — close to, but not quite, the one tenth of the price the company claims. A 50 percent promotion runs until 9 September 2026, which for now doubles the gap again. The performance claims stay in the maker’s own hands. Z.ai says the model beats GLM-5.2 across benchmarks and approaches Claude Opus 4.8 on coding and agentic tests, but publishes those results only as a chart image, and no independent laboratory has repeated them. What can be verified is the licence: the LICENSE file in the repository is the bare MIT text, with no attribution rider and no restriction on commercial use.
GLM-5.3-Flash →OpenAI still calls a February model its best coder — and it knows nothing written after August 2025
OpenAI's documentation describes GPT-5.3-Codex as "the most capable agentic coding model to date". It went on sale on 24 February 2026. Six months later that sentence still stands, because nothing has replaced it: the dedicated Codex line has not had a new member since. The general line has not stood still in the same period. GPT-5.5 arrived on 24 April and the three-model GPT-5.6 family — Sol, Terra and Luna — on 9 July, and OpenAI's own catalogue entry for GPT-5.6 Sol now reads "start here for complex reasoning and coding". The specialist and the generalist are being pointed at the same job. The gap between them is not marketing. GPT-5.6 Sol carries a 1,050,000-token context window against the Codex model's 400,000, and a knowledge cutoff of 16 February 2026 against 31 August 2025. For a coding model the second number matters more than it looks: a cutoff of August 2025 means the tool writing your dependency call has never seen a year of releases, deprecations and breaking changes in the libraries it is calling. What the older model keeps is price. GPT-5.3-Codex costs $1.75 per million input tokens and $14 per million output, against $4 and $20 for GPT-5.6 Sol — a bit over half the input rate. It is also narrower by design: the Codex line runs only on the Responses API, with no Chat Completions, no batch, no fine-tuning. That leaves a choice OpenAI does not spell out anywhere in one place. The cheap specialist is frozen in time; the expensive generalist is current. Nothing in the documentation says the Codex line has been retired, and no shutdown date has been published for it — which is exactly why the six-month silence is worth noticing rather than assuming. Figures in this article come from OpenAI's own model cards and pricing tables, read on 26 August 2026; the release dates were cross-checked against an independent model registry, because OpenAI does not date its model cards.
OpenAI is switching off its entire image line by 1 December — and the cheapest model dies with it
OpenAI's deprecation page now schedules the end of every image model the company sells except one. gpt-image-1 goes dark on 23 October 2026; gpt-image-1.5, gpt-image-1-mini and chatgpt-image-latest all follow on 1 December. Every one of them names the same replacement: gpt-image-2. For most users that is a price cut. Priced per million image tokens, the 2025 original gpt-image-1 is the dearest model in the whole line at $10 in and $40 out. gpt-image-1.5 charges $8 and $32. The survivor, gpt-image-2, charges $8 and $30 — 25% less on output than the model it replaces, which is the opposite of what the industry has been doing this year. The exception is the one model built for people watching costs. gpt-image-1-mini charges $2.50 per million image tokens in and $8.00 out. It has no successor of its own: the migration table sends it to gpt-image-2 as well. That is 3.2 times more on input and 3.75 times more on output, for anyone who chose the mini precisely because it was cheap. The pattern is becoming familiar. Cheap tiers are announced as an entry point, then quietly folded into the flagship when the line is consolidated — and the bill for the migration lands on the users who were most price-sensitive to begin with. Batch pricing halves all of these figures, but it halves them for both the old model and the new one, so the ratio does not move.
Mistral retires Medium 3 in eight days — the replacement it names costs 3.75 times more per token
On 31 August 2026 Mistral AI switches off two models at once: Mistral Medium 3 (API name mistral-medium-2505, released May 2025) and Mistral Medium 3.1 (mistral-medium-2508, August 2025). Both were marked deprecated on 22 May 2026, and both are billed at $0.40 per million input tokens and $2.00 per million output. The company's documentation names a single replacement for them: Mistral Medium 3.5. That replacement is listed at $1.50 input and $7.50 output. The multiplier is the same in both directions — 3.75 times. A team that keeps its prompts, its volumes and its code exactly as they are, and does nothing but follow the migration path printed in Mistral's own table, will see its bill for this model line grow by 275% on the first of September. The part that is easy to miss is that the cheaper option is upwards, not sideways. Mistral Large 3, the company's larger and older model from December 2025, costs $0.50 input and $1.50 output — a third of Medium 3.5's input price and a fifth of its output price. On Mistral's public price list the ordering is inverted: the mid-tier model is the expensive one, and the model above it in the naming scheme is cheaper than both the model it sits above and the model it replaces. Below them, Mistral Small 4 runs at $0.15 and $0.60. The inversion is worth noting because of how Medium 3 was sold. Mistral launched it in May 2025 under the headline "Medium is the new large", with a pitch built almost entirely on price: roughly eight times cheaper than comparable models, deployable on four GPUs, at or above 90% of Claude Sonnet 3.7 on the company's own benchmarks. Fifteen months later the tier that was created to be the cheap one is the one that costs the most per token in the range. None of this makes Medium 3.5 a bad model — it is newer, and Mistral positions it as the stronger of the two. But the migration note in the documentation says only "use Mistral Medium 3.5 for new integrations", and says nothing about what that costs. Anyone still calling mistral-medium-2505 or mistral-medium-2508 has eight days to decide whether the named successor, the larger model or the smaller one is the right destination.
Mistral Medium 3 →xAI's new video model costs 60% more per second — and is the one model its own batch discount refuses
xAI's developer documentation now lists two generations of its video model side by side, and the gap between them is not only technical. Grok Imagine Video 1.5, whose full set of modes was announced on 31 July 2026, is billed at $0.080 per second of generated footage. The first-generation grok-imagine-video remains on the price list at $0.050 per second — a 60 percent difference for the newer model. What the higher rate buys is a native 1080p pipeline for text-to-video and image-to-video, rather than a lower-resolution render enlarged afterwards, plus a reference mode that guides a clip with several images without locking the opening frame, and the option to attach up to three preset voices to the subject. Clips run from one to fifteen seconds; separate endpoints extend an existing clip from its last frame or edit one, the latter capped at 720p and about 8.7 seconds. The less advertised part sits in the batch documentation. xAI's Batch API, which processes large volumes asynchronously at a discount, accepts image and video requests only for the first-generation grok-imagine-image and grok-imagine-video models. Both 1.5-generation models — the video model and Grok Imagine Image 2.0 — are turned away with an explicit "not supported for batch processing" error. For anyone generating footage in bulk, the newer model is therefore dearer twice over: a higher per-second rate, and no access to the cheaper queue. The documentation also discloses how text-to-video actually works on this model. Rather than a single pass, xAI writes, the model generates a first frame from the prompt and then animates it; the intermediate image is never returned to the caller, though it is a single billable request.
Grok Imagine Video 1.5 →ByteDance's cloud switched on the DeepSeek price rise — and left the older build of the same model untouched at a quarter of the price
The increase BytePlus announced five days ago is now live. Since 00:00 Beijing time on 21 August 2026, ModelArk bills deepseek-v4-flash-ga-260731 at $0.44 per million input tokens, $1.32 per million output tokens and $0.014 for cache hits. Batch inference is charged at exactly half: $0.22 and $0.66. What the notice does not mention is the line directly below it in the same table. The April build of the same model, deepseek-v4-flash-260425, is still on sale and its rates were not touched: $0.14 in, $0.28 out. That is 3.1x less on input and 4.7x less on output than the version the platform now recommends. Only one number moves the other way — cache-hit input on the old build costs $0.028, twice the new rate. The rise puts ModelArk exactly level with DeepSeek's own list price, but only at one hour of the day. The vendor charges $0.44 / $1.32 during peak hours (01–04 and 06–10 UTC) and precisely half of that for the remaining eighteen hours. ModelArk has no time-of-day rate at all. The only way to halve the bill there is batch inference — the same 50% discount, paid for with latency instead of a clock. The parity is not limited to the small model. deepseek-v4-pro-ga-260813, the final August build of the larger model, is listed on ModelArk at $1.32 input and $3.96 output — again the vendor's peak rate to the cent, again with no off-peak equivalent.
DeepSeek-V4-Flash →Every pinnable ChatGPT model has now been switched off in OpenAI's API — what is left costs four times more
OpenAI has always sold two different things under similar names: the models in its API, and the model that actually answers inside ChatGPT. The second kind was reachable through a small family of aliases — chatgpt-4o-latest, then gpt-5-chat-latest and its successors. As of 10 August 2026, every one of them that carried a version number has been switched off. The sequence is short. chatgpt-4o-latest was removed on 17 February 2026. gpt-5-chat-latest and gpt-5.1-chat-latest were shut down together on 23 July 2026, under an announcement made on 22 April. gpt-5.2-chat-latest and gpt-5.3-chat-latest followed on 10 August 2026, deprecated on 8 May. In each case OpenAI named GPT-5.6 Sol — an API model, not a chat one — as the replacement. What remains is a single entry called chat-latest, and the wording on its card is the point: it "points to the latest Instant model currently used in ChatGPT" and "the underlying model snapshot will be regularly updated". There is no version number to pin. A developer who wants the behaviour of the ChatGPT model can still have it, but not frozen, and not with a guarantee that today's answers will resemble next month's. The price moved in the same direction. gpt-5-chat-latest and gpt-5.1-chat-latest cost USD 1.25 per million input tokens and USD 10 per million output. The 5.2 and 5.3 snapshots raised that to USD 1.75 and USD 14. chat-latest is listed at USD 5 and USD 30 — four times the input price of the first two, and precisely the price of the GPT-5.6 Sol flagship. Part of that buys real capability: the window grew from 128,000 tokens to 400,000, and the output ceiling from 16,384 tokens to 128,000, an eightfold increase. But the pinnable versions were the cheap ones, and they are the ones that are gone. There is also a rule behind the timing, and OpenAI publishes it. Its deprecation policy promises generally available models at least six months of notice, and specialised variants at least three — with "chat variants such as gpt-5.1-chat-latest" given as the first example of the latter. The line closest to the consumer product carries the shortest guarantee in the catalogue. One footnote for anyone reading the documentation directly: the card for GPT-5.1 Chat still describes it as the snapshot currently used in ChatGPT and invites developers to test chat improvements on it. The deprecation table has overtaken that sentence. And a distinction worth keeping straight — what ended was the API alias, not the model's presence in ChatGPT itself.
OpenAI Chat Latest →OpenAI's voice models have had one price cut in 22 months — and text output got dearer along the way
OpenAI has told developers that nine legacy voice models — the entire GPT-4o audio and realtime line, plus the first generation of GPT-Audio and GPT-Realtime — stop answering on 20 January 2027. Adding the four missing GPT-4o voice profiles to our catalogue put the whole price history of that line in one place, and it reads differently from the price history of text. When GPT-4o Audio and GPT-4o Realtime went live on 1 October 2024, an audio token cost $40 per million on the way in and $80 on the way out. In August 2025 the successors, GPT-Audio and GPT-Realtime, brought that to $32 and $64. Today's generally available replacements — gpt-audio-1.5 and gpt-realtime-2.1, the models OpenAI names in the shutdown notice — charge exactly the same $32 and $64. That is one cut, of 20%, in twenty-two months, and nothing since. Text inside the same models moved in both directions. In the realtime line, input fell from $5 to $4 per million and cached input collapsed from $2.50 to $0.40, a drop of 84%. Output went the other way: $20 in 2024, $24 today — 20% more expensive. And the ratio that matters for anyone building a voice product has not shifted at all. In October 2024 an audio input token cost eight times a text input token; in August 2026, at $32 against $4, it still costs eight times as much. What the buyer does get for the same money is room. GPT-4o Realtime remembered 32,000 tokens and answered with at most 4,096; gpt-realtime-2.1 holds 128,000 and returns up to 32,000 — four times the context and eight times the answer at an unchanged audio rate. The cheap tier tells the same story from below: GPT-4o mini Realtime, launched on 17 December 2024, had the shortest memory in the family at 16,000 tokens. One caveat: these are list prices from OpenAI's own model cards, not what any particular customer pays, and audio tokens are not directly comparable to text tokens — a second of speech consumes far more of them than a second of reading. The comparison here is strictly OpenAI against itself, one voice generation against the next.
GPT-4o Realtime →DeepSeek's clock is a Beijing clock: the same working day costs 75% more there than in San Francisco
DeepSeek's price list, in force since 16 August 2026, is the first from a major model provider in which the hour of the day is a full pricing variable rather than a promotion: peak hours are 01:00–04:00 and 06:00–10:00 UTC, and everything outside them costs exactly half. The published windows say nothing about where those hours fall. Converted, they say a great deal. Beijing keeps UTC+8 all year, with no daylight saving. The two peak windows land at 09:00–12:00 and 14:00–18:00 local time — Chinese office hours. The gap between them, 04:00–06:00 UTC, is 12:00–14:00 in Beijing: the lunch break, priced at half rate. The tariff is not described in geographic terms anywhere in the documentation, but it is drawn around one country's working day. That has an arithmetic consequence nobody has published. Take a team using DeepSeek-V4-Pro evenly through a 09:00–17:00 local working day, and price a million output tokens ($3.96 at peak, $1.98 off-peak). In San Francisco the local working day is 16:00–24:00 UTC and misses both peak windows entirely, so every hour bills at $1.98. In Warsaw it is 07:00–15:00 UTC, of which three hours are peak: $2.72 per million, 37.5% above San Francisco. In Beijing it is 01:00–09:00 UTC, of which six hours are peak: $3.47 per million — 75% above San Francisco, for identical work on identical weights. Bangalore lands in between at $3.09, or 56% above. European bills also move with the clock change. In winter Warsaw shifts to UTC+1, only two working hours stay inside the peak, and the same million falls to $2.48 — a 9% discount granted by nothing but the end of daylight saving. What this does not show: DeepSeek is not charging anyone by location. The tariff is identical worldwide and time-based, and the company presents it as load management — the peak windows are simply when its servers are busiest, which is when its home market is at work. Batch and overnight jobs can be moved into the cheap hours by anyone, anywhere, and for most production workloads cached input, billed at a fiftieth of a cache miss, matters far more than the clock. The figures above assume usage spread evenly across office hours, which no real team does exactly. But the direction is not an artefact: the further a user's working day sits from Beijing's, the less DeepSeek's price rise costs them.
DeepSeek-V4-Pro →Amazon's newer voice model listens 12% cheaper — and writes 11 times dearer
Amazon never published a price comparison between its two speech-to-speech models, but the Bedrock price list does it for anyone who reads all four rates instead of the headline one. In the us-east-1 list dated 13 August 2026, Nova 2 Sonic charges $3.00 per million speech input tokens and $12.00 per million speech output, against $3.40 and $13.60 for the original Nova Sonic. That is a cut of 11.8% on both sides of the audio. The text rates moved the other way, and much harder. Text input went from $0.06 to $0.33 per million tokens — five and a half times more — and text output from $0.24 to $2.75, a rise of 11.5 times. Both models bill audio and text separately, so a session's bill depends on how the two mix. That mix has an arithmetic break-even. Moving to Nova 2 Sonic saves $1.60 per million speech output tokens and adds $2.51 per million text output tokens. The newer model is therefore cheaper only while text output stays below roughly 64% of speech output volume. A model that mostly talks stays cheap; one that also returns transcripts, tool calls and structured replies crosses the line. What this does not prove: we have not measured how either model actually splits its tokens in a real session, and Amazon publishes no such breakdown, so the 64% figure is a property of the price sheet, not a verdict on any application. Regional lists other than us-east-1 may differ, and both sets of rates can change without notice.
Amazon Nova 2 Sonic →Kimi's platform drops seven model IDs on 31 August — and the flagship it points to charges 3× more per output token
Moonshot AI is retiring its entire classic generation line. The model list on the Kimi API platform now carries a single note: following the Kimi K3 launch, `kimi-k2.5` and the `moonshot-v1` series are no longer available to newly registered users, with a full platform sunset on 31 August. That covers seven model identifiers — moonshot-v1 in 8k, 32k and 128k variants, the three matching vision-preview versions, and Kimi K2.5, the model the company itself described as open-source state of the art for agentic, code and vision tasks. Four identifiers survive: K3, K2.7 Code, K2.7 Code Highspeed and K2.6. The cost of moving up is on the same platform's price pages. The largest of the departing models, moonshot-v1-128k, is billed at $2.00 per million input tokens and $5.00 per million output, with a 131,072-token context window. Kimi K3 is billed at $3.00 per million input tokens on a cache miss, $0.30 on a cache hit, and $15.00 per million output — three times the output rate and one and a half times the input rate, for a context window eight times larger at 1,048,576 tokens. Users on the cheapest departing model, moonshot-v1-8k at $0.20 and $2.00, face a steeper jump: fifteen times the input rate and seven and a half times the output rate. What makes this worth recording is the cadence, which only becomes visible when the dates on that one page are read together. Moonshot has switched off kimi-thinking-preview on 11 November 2025, kimi-latest on 28 January 2026, the five-model kimi-k2 series on 25 May 2026, and now the V1 line and K2.5 on 31 August. The gaps are 78, 117 and 98 days: on average, this vendor removes a model family roughly every three months. Self-hosting is the escape hatch that a closed platform does not offer — K3's weights are published under a modified MIT licence with a commercial threshold, and K2.5 was an open-weights release too, so anyone running them on their own hardware is unaffected by the API shutdown. One caveat on the date: the documentation gives only the day and month, without a year. The note is framed as a consequence of the K3 launch, which took place on 16 July 2026, so the deadline falls on 31 August 2026 — twelve days from now. Moonshot published no separate deprecation announcement; the dates live in a table inside the developer documentation.
Kimi K3 →OpenAI's price ladder for its flagship doubles at every rung — and the only tier with a published speed is the one you cannot buy
Five days after OpenAI published the first tokens-per-second figure in the history of its API — 750 output tokens per second for the Ultrafast preview of GPT-5.6 Sol — that tier still has no price. The tier that does have a price still has no speed. The two facts belong together, and the shape of OpenAI's price list makes the point sharper than either announcement does. OpenAI sells GPT-5.6 Sol at three priced service levels, and on its own pricing page each level is exactly double the one below it. Flex and Batch processing cost USD 2.50 per million input tokens and USD 15 per million output. Standard costs USD 5 and USD 30. Fast mode costs USD 10 and USD 60. The doubling also holds on the long-context rows, which apply to requests above the model's 272,000-token threshold: 5 and 22.50 for Flex, 10 and 45 for Standard, 20 and 90 for Fast. Eight numbers, four exact factors of two, no rounding and no exceptions. What the ladder does not carry is a single speed. Fast mode carries a 100% premium over Standard, and OpenAI has never published a tokens-per-second rate, a latency figure or a percentage to say what that premium delivers. Flex is documented only as offering "slower response times and occasional resource unavailability" — again with no number attached. A customer choosing between the rungs is choosing between prices that are precise to the cent and speeds that are not stated at all. Ultrafast inverts this exactly. It is the first service level OpenAI has attached a throughput figure to — "up to 750 output tokens per second", "up to 14× faster than Standard processing" — and it is the only one with no price, available in limited preview to a select group of customers. One baseline can be derived from those two figures, with a caveat that matters. If the 750 tokens per second and the 14× multiplier describe the same run, Standard processing generates roughly 54 output tokens per second. That would be the closest thing to a published baseline this model has. But both figures are ceilings marked "up to", and OpenAI does not state that they were measured on the same run, so 54 is an inference from the announcement rather than a number OpenAI has disclosed. What this does not prove: nothing here indicates what Ultrafast will cost when it is priced. The doubling across the three existing rungs describes the current list, not a commitment about the next one, and a preview price need not survive to general availability.
GPT-5.6 Sol →The cheapest lab stops being cheap: DeepSeek raises API prices from 16 August and starts charging by the clock
DeepSeek has published a new price list for its V4 models, effective 16 August 2026. The company that built its reputation on undercutting everyone else is raising rates across the board — by between 50 percent and more than 1,100 percent, depending on the model, the type of token and the hour of the day. The numbers for DeepSeek-V4-Flash: output tokens go from $0.28 to $1.32 per million at peak, cache-miss input from $0.14 to $0.44, and cached input from $0.0028 to $0.014 — a fivefold rise on the cheapest line in the catalogue. DeepSeek-V4-Pro goes from $0.87 to $3.96 per million output tokens at peak, while its cached input rises from roughly $0.0036 to $0.044 per million: the 1,100 percent figure comes from that one line, not from the headline rate. The structural change matters as much as the numbers. From Sunday the price depends on when the request is made: peak hours are 01:00–04:00 and 06:00–10:00 UTC, everything else is off-peak and costs exactly half. DeepSeek says the tiered structure is meant to "allocate resources more reasonably" and push developer workloads towards less congested windows. In practice it is the first time a major model provider has made the hour of the day a first-class pricing variable rather than a promotional discount. The timing is what makes this striking. In the same week Anthropic cancelled a planned 50 percent rise for Claude Sonnet 5 and Google put an expiry date on its Gemini 3.7 Flash discount, DeepSeek moved in the opposite direction — and moved further than either. Off-peak V4-Flash output at $0.66 per million is still cheap by Western standards, but the gap that made DeepSeek an obvious default has narrowed by a factor of four or five. Both V4 models keep their 1M-token context window and 384K maximum output; the concurrency limits (2,500 simultaneous requests for Flash, 500 for Pro) are unchanged. We have updated the pricing in both profiles in the catalogue.
DeepSeek-V4-Flash →The 50% rise is off: Anthropic makes Claude Sonnet 5's $2/$10 the standard price
Sixteen days before the deadline, Anthropic has cancelled the price increase it scheduled for Claude Sonnet 5. The company's pricing documentation now states plainly that the $2 per million input tokens and $10 per million output tokens, announced at launch as an introductory rate running through 31 August 2026, is the standard price, and that the increase to $3 and $15 planned for 1 September will not occur. This reverses the situation we reported on 6 August, when the same documentation carried the expiry date as a footnote. What was a discount with a countdown is now simply the price. Sonnet 5 keeps its position in Anthropic's ladder — Fable 5 at $10 / $50, Opus 5 at $5 / $25, Sonnet 5 at $2 / $10, Haiku 4.5 at $1 / $5 — but it is no longer the only current Claude model sold below its own list. One qualification from that earlier piece still holds, because it was never about the rate card. Sonnet 5 uses the tokenizer introduced with Opus 4.7, and Anthropic's own note says the same text yields roughly 30% more tokens on it than on the previous generation; Sonnet 4.6 and earlier use the older tokenizer. On the headline, Sonnet 5 is a third cheaper than Sonnet 4.6's $3 / $15. Measured per page of text rather than per token, the saving is closer to a tenth. The same arithmetic applies to the context window: the company puts one million tokens on Sonnet 5 at about 555,000 words, against roughly 750,000 words for the same million on Sonnet 4.6. The timing invites one comparison. Two days ago Google presented Gemini 3.7 Flash at $0.75 and $3.75 per million tokens, described as half the price of its predecessor — a rate its own price list dates to 31 December 2026, after which it doubles. Within one week, then, two vendors have taken opposite decisions about the same instrument: one has removed the expiry date from a discount, the other has left it in place. For anyone budgeting a year ahead, that difference matters more than the headline figures. Our Claude Sonnet 5 profile has been updated with the standard price and the cancelled increase.
Claude Sonnet 5 →Claude Sonnet 5's introductory price expires on 31 August — and the bill rises 50%
**Update, 15 August 2026:** Anthropic has cancelled this increase. Its pricing documentation now names $2 / $10 the standard price for Claude Sonnet 5 and states that the rise to $3 / $15 on 1 September will not occur. See: The 50% rise is off (/news/claude-sonnet-5-price-rise-cancelled-2-10-permanent). The paragraph below on the tokenizer still applies. Anthropic's model documentation carries a footnote that is easy to miss and expensive to ignore: the introductory pricing of $2 per million input tokens and $10 per million output tokens applies to Claude Sonnet 5 only through 31 August 2026. From 1 September the model reverts to its list price of $3 and $15 — a 50% increase on both sides of the meter, arriving in under four weeks. Sonnet 5 is not a marginal product for Anthropic. Released on 30 June 2026, it is the default model for Claude's free and Pro users and the company's declared best combination of speed and intelligence, with a native one-million-token context window, 128,000 output tokens on the Messages API and 63.2% on SWE-Bench Pro. It is also, as of today, the only current Claude model sold below its own list price: Fable 5 stands at $10 / $50, Opus 5 at $5 / $25 and Haiku 4.5 at $1 / $5, all at list. There is a second, quieter movement in the same direction, and it has nothing to do with the price card. Anthropic's own documentation puts the one-million-token window of Fable 5, Opus 5 and Sonnet 5 at roughly 555,000 words, while the same one million tokens on Opus 4.6 and Sonnet 4.6 held about 750,000 words. For Fable 5 the company states the mechanism outright: it uses the tokenizer introduced with Opus 4.7, and the same text produces roughly 30% more tokens than on the older models. Since tokens are the billing unit, an unchanged workload on the newer generation is metered higher before any rate change is applied at all. What this does not appear to be is a repricing of the frontier. Anthropic has not announced a change to any other model's rates, and the retirement floor for Sonnet 5 remains unmoved at 30 June 2027. The plain reading is that a launch discount is simply running out on schedule. For anyone budgeting against Sonnet 5, the two dates that matter are 31 August, when the discount ends, and the moment a workload is ported from a 4.6-generation model, when the token count itself changes. Our Claude Sonnet 5 profile now carries both the introductory and the list price.
Claude Sonnet 5 →