News
What's happening in robotics and AI — curated by the wujec.ai editors.
Anthropic's price list sells Claude Haiku 3.5 on two clouds. Both switched it off months ago
This catalogue recently left a loose end hanging. Anthropic's price list carries Claude Haiku 3.5 as retired everywhere except Amazon Bedrock and Google Cloud, while Google Cloud had published a shutdown date of 5 July 2026 for that same model. One of the two documents had to be out of date. So we went and read the clouds themselves. Amazon Bedrock's model card for Claude 3.5 Haiku, read on 16 September 2026, gives a model EOL date of 19 June 2026. Bedrock's own lifecycle rules state what EOL means there: after that date the model is removed from all AWS Regions and requests made to it fail. Google Cloud's partner-model deprecation page gives 5 July 2026 — sixteen days later, and also in the past. Anthropic's own API stopped answering on 19 February 2026. That is every place the model ran. There is no fourth. The price list, checked the same day, still names both clouds, and it is not a stub entry. The model keeps a full rate card: base input, both cache-write tiers, cache hits, output, and separate batch pricing. A developer reading that page today would reasonably conclude the model can still be bought from a partner. It cannot. Because a single stale line proves little, we checked the other three models flagged the same way. Claude Sonnet 4 holds up: Bedrock lists it in the Legacy state with an EOL date of 14 October 2026, and it entered paid extended access on 14 July. Claude Opus 4.1 holds up: Legacy, EOL 8 January 2027. Claude Opus 4 is flagged as retired except on Google Cloud, and that narrower claim is also right — Google Cloud still shows it as generally available, while Bedrock's catalogue no longer carries a model card for it at all. One claim in four is wrong, and it concerns the oldest and cheapest model of the group. The pattern is worth naming, because it cuts both ways. A vendor's price list describes what the vendor intends to sell; the partner clouds decide, on their own calendars, what they actually serve. Neither document is authoritative about the other. For a model that has left the producer's own API, the availability line in a price list is the weakest evidence available, and the partner's lifecycle page is the strongest. One correction of our own belongs here. Our profile of Claude 3 Haiku, the 2024 model, said it fell silent everywhere on 23 August 2026, when Google Cloud shut it down. Bedrock kept serving it to an EOL date of 10 September 2026 — eighteen days longer, and only six days before this article. The profile has been fixed, along with the four others named above.
Claude 3.5 Haiku →Anthropic switched four models off its own API and kept their prices. They cost three times the model that replaced them
A model can be switched off in two different places, and Anthropic's own documents show how far apart those places can be. The company's deprecation page is unambiguous: claude-opus-4-20250514 and claude-sonnet-4-20250514 were retired on 15 June 2026, claude-opus-4-1-20250805 on 5 August 2026, and claude-3-5-haiku-20241022 back on 19 February 2026. On the Claude API they cannot be called at all. The price list, checked on 22 September 2026, still carries all four. Each one is labelled — Opus 4.1, Sonnet 4 and Haiku 3.5 as retired “except on Bedrock and Google Cloud”, Opus 4 as retired “except on Google Cloud” — and each keeps a full set of rates: base input, both cache write tiers, cache hits, output, and separate batch pricing. These are not historical footnotes. They are the prices a customer pays today, through a partner cloud. What that produces is an inverted ladder. Claude Opus 4 and Opus 4.1 are billed at $15 per million input tokens and $75 per million output. Claude Opus 5, Opus 4.8, 4.7, 4.6 and 4.5 — every newer model in the line — are billed at $5 and $25. The two models Anthropic no longer serves itself are three times the price of the five it does, and the gap widens on batch work, where the retired pair costs $7.50 per million input against $2.50 for the current generation. The explanation is mundane: those rates were set in 2025 and never revised, because for Anthropic the products are finished. Nothing forces a price cut on a model you have stopped selling. But the consequence is not mundane for anyone still running production traffic on Opus 4.1 in Bedrock or Vertex: the upgrade path is not a cost, it is a two-thirds discount. One detail does not line up. Anthropic's price list flags Haiku 3.5 as still available on Google Cloud, while Google Cloud published a shutdown date of 5 July 2026 for that model. On a page the vendor maintains itself, the partner availability note has apparently outlived the partner's own schedule — one more reason to read a price list as a statement of intent rather than a live inventory.
Claude Opus 4.1 →xAI stops charging for searches on X and starts charging for what they bring back — including posts nobody asked for
From 21 September 2026 at 12:00 PT, xAI bills its X Search tool by the item it returns rather than by the request that triggered it. The documented rate changes from 5 US dollars per 1,000 calls to 5 US dollars per 1,000 posts fetched and 10 US dollars per 1,000 user profiles fetched. Under the old rule the size of an answer did not matter: one call cost half a cent whether it came back with one post or fifty. Under the new one the same call costs half a cent for every post in the response, and a search over user accounts costs a full cent for each profile returned. The documentation states the counting rule plainly: "Every post returned by a search or thread fetch, including parent and quoted posts, counts", and "every profile returned by a user search counts". The clause about parent and quoted posts is the part worth reading twice. When the tool pulls a thread, the posts above the one being examined and the posts it quotes are billed as well, even though nothing in the query asked for them. The developer does not choose them; the structure of the conversation on X does. Neither the pricing page nor the tool documentation names a ceiling for how many posts a single call may return, and no parameter for capping that number appears in the published description of the tool. The upper bound of what one search can cost is therefore not stated anywhere in the documentation — a change of kind, not only of rate, because the old model had a fixed maximum per call by construction. X Search is now the only server-side tool xAI prices this way. Web search and code execution remain at 5 dollars per 1,000 calls, file attachment search at 10 dollars and collections search at 2.50 dollars per 1,000 calls — all counted per invocation. xAI's own guidance notes that the agent decides how many tools to call, so "costs scale with query complexity"; for searches on X, cost now scales with how talkative the platform is as well.
Grok 4.6 →OpenAI sells three cybersecurity models. Two of them exist only as a row in the price list
OpenAI's GPT Cyber line — models trained for authorised vulnerability research and sold only to defenders vetted through the company's Daybreak programme — has had at least three versions in 2026. Only one of them has a model card. The public model catalogue at platform.openai.com lists exactly one Cyber model: GPT-5.6 Cyber, launched on 10 August 2026 with a full specification, benchmark comparisons and an endpoint table. The documentation addresses for gpt-5.5-cyber and gpt-5.4-cyber return 404. Neither model appears in the company's changelog, which otherwise records API releases back to 2023. The price list tells a different story. It currently carries a row for gpt-5.5-cyber at USD 12.50 per million input tokens and USD 75 per million output — the same rates as the documented GPT-5.6 Cyber. Internet Archive snapshots of that page show both older models appearing between 3 May and 5 June 2026: GPT-5.5 Cyber at USD 20 input and USD 120 output, cut by 1 August to today's level, a reduction of 37.5 percent. The row for GPT-5.4 Cyber showed dashes in every column throughout, so its price was never published at all. GPT-5.4 Cyber has now left the price list as well. OpenAI's deprecation page says it is removed from the API on 1 October 2026, with GPT-5.6 Cyber as the migration target — an announcement made on 11 September, twenty days ahead, where the rest of the company's 2026 schedule allows roughly six months. The one capability figure published for any of these models is comparative and came with the successor's launch: on OpenAI's internal Advanced Cybersecurity Completion Rate, which counts how often a model answers dual-use security prompts instead of refusing, GPT-5.5 Cyber reaches 57.3 percent against 95.0 percent for GPT-5.6 Cyber. No test of actual skill has been released for either older version. The gap matters because the Daybreak Red alias points at the newest model only. A customer running GPT-5.5 Cyber has to name it explicitly — and has no manufacturer documentation describing what they are running.
GPT-5.5 Cyber →AWS took the promise that a Bedrock model lives a year out of the rulebook — and put it in the model card
Amazon has split the rules that decide how long a model on Amazon Bedrock is guaranteed to answer. Models launched before 7 September 2026 keep the old policy; everything launched on or after that date falls under a new one, and the new one no longer contains the two sentences customers used to plan around. The old page — still live, now titled "Model lifecycle (Legacy)" — makes two platform-wide promises. First: "Once a model launches on Amazon Bedrock, it will remain on Amazon Bedrock for at least 12 months before the EOL date." Second: "A model will be in the Legacy state for at least 6 months before the EOL date" — that is, at least half a year of notice before the switch-off. Neither sentence appears on the new page. What appears instead is a per-model arrangement. The new policy says every model card shows two things: an "EOL no sooner than" date, and the Legacy period, meaning the notice given before end of life. And the notice period is no longer one number: "There are two Legacy periods: 6 months and 45 days. Most models have a 6-month Legacy period." A 45-day wind-down is now written into the rules as an option a provider can take. So far nobody has taken it. We read nine model cards across Amazon, OpenAI, Anthropic, xAI, Z.ai, MiniMax, Moonshot and TwelveLabs, including the newest arrival on the platform: GPT-6 Astra, launched 8 September 2026 — one day after the cutover, and therefore among the first models governed by the new policy. Its card reads "EOL no sooner than: September 8, 2027" and "Legacy period: at least 6 months". Twelve months of life, six months of notice: exactly the old terms, now stated for one model rather than promised to everyone. One clause went the customer's way. The old policy describes a "public extended access" phase: after at least three months in Legacy, a model enters a stretch in which it keeps running but, as AWS puts it, "you should expect higher pricing, which will be set by the model provider". The new policy does not mention public extended access at all. The last months of a model's life were the leg where the bill could rise; that leg is not in the new rules. The practical consequence is small today and structural tomorrow. Anyone building on a hosted model has to answer one architectural question — how long will this thing keep answering — and until now Bedrock answered it once, for the whole catalogue. Now the answer sits in each model card, set per model, and a provider who wants to move fast has a 45-day door available. The number to check is no longer in the terms; it is on the card. The old rules, meanwhile, are still doing their work. On 14 September 2026 Bedrock switched off Amazon Nova Sonic, the company's first speech-to-speech model, and Amazon Nova Premier, its first frontier-class model and the designated teacher for the cheaper Nova line. Both were moved to Legacy on 13 March 2026 — six months' notice to the day. Nova Canvas and Nova Reel, Amazon's image and video models, go on 30 September, and Amazon has named no successor for either.
Amazon Nova Premier →Inception sells new customers exactly one Mercury model — its own documentation still prices the three it closed
A footnote at the bottom of Inception's models page settles what the company is actually selling today: "Mercury 1, 2, and Mercury Edit 2 remain supported for existing customers." Access and migration, it adds, go through an Inception representative. For anyone opening an account this week, that leaves one buyable model — Mercury 2.5, announced on 8 September — plus two previews, Mercury Voice and Mercury Router, whose pricing is a request to contact sales. The interesting part is that the company's own developer documentation has not caught up with its marketing. The models reference page still carries Mercury 2 and Mercury Edit 2 in the same table as the flagship, with full rates ($0.25 per million input tokens, $0.75 output, $0.025 cached), their context windows (128K for chat, 32K on both code endpoints) and their endpoints. The FAQ on the same site answers a question about image input by naming all three models in the present tense. Nothing there tells a developer that two of them can no longer be bought. That matters for a catalogue like this one, because a live model card is not the same thing as a live product. Both models are running and billable — just not to new accounts. Neither has a published shutdown date, which is why our profiles keep them in production rather than marking them retired. The timeline is worth stating plainly. Mercury 2 was announced on 24 February 2026 as, in Inception's words, the world's fastest reasoning language model, at 1,009 tokens per second on NVIDIA Blackwell cards. Six and a half months later its successor arrived with 260K of context instead of 128K and a list price of $0.20 input — cut to $0.04 by a launch discount — against Mercury 2's unchanged $0.25. Mercury 2 was not undercut by a rival. It was undercut by the next model from the same company, and then quietly taken off the shelf. For a company whose entire pitch is speed, the product line moves at a similar pace. Buyers of a diffusion model should read that footnote as part of the price.
Mercury 2 →A 300 kg robot you ride like a horse is not the surprise. The surprise is that it is sold in an online shop
DAX Robotics showed its Qiji rideable quadrupeds at the World Robot Conference in Beijing in August 2026, and the pictures travelled further than the facts. The machine that matters is Qiji X1: 300 kg, sixteen driven joints fed from a 165 V pack, a peak joint torque of 1400 Nm, 8 to 10 hours of work or roughly 40 km per charge, at a walking pace of 7 to 10 km/h. The rider sits in a saddle and steers with handlebars; balance and footing are the robot's problem, not the rider's. The number worth pausing on is the payload: 300 kg dynamic load, exactly what the robot itself weighs. A machine that carries its own mass is not a scaled-up robot dog, and DAX does not pretend it is a road vehicle either — with point feet and no wheels, the X1 is aimed at slopes, gravel, mud and unpaved trails. The faster wheel-legged sibling, Qiji XS, is quoted at 320 kg, over 40 km/h and 60 km of range, and the heaviest member of the line, the T1000 presented in Beijing in April 2026, is rated for a tonne of cargo on joints delivering more than 2000 Nm. The genuinely new thing is the sales channel. The series goes out under a three-year exclusive arrangement with JD.com — a consumer marketplace, with early-buyer sweeteners of the kind used for phones: discounts, five years of free servicing, free over-the-air updates. Heavy legged machines have until now been sold the way industrial equipment is sold, through integrators and quotations. Putting one on a retail platform assumes there is a private buyer at the other end, and reservations reported in July at over three thousand suggest the assumption is being tested rather than proven. One caution for readers who saw this robot elsewhere. Much of the international coverage put the machine at 180 kg and dated the launch to 8 September; the manufacturer's figure is 300 kg, and the Beijing show ran in August. The prices circulating — around 289,000 yuan — come from press reports, not from a DAX price list, and the company has published no evidence of series deliveries. In this catalogue the robot is therefore listed as a prototype, however open the order book may be.
Qiji X1 →The humanoid Japan chose for mass production has 19 joints and lifts 3 kg — and its brain is still being built
Mitsubishi Motors will put a humanoid robot on an unused engine line at its Kyoto plant, with production targeted for early 2027. The machine is HL Human, built by Highlanders, a University of Tokyo spin-off founded in May 2023. Read the specification sheet next to the announcement and the two do not obviously match. HL Human, unveiled on 30 June 2025, has 19 degrees of freedom — four per arm, the rest in legs and torso — and grips objects of up to 3 kg, with a five-finger hand offered as an option. For comparison, the humanoids shipping out of China in volume today carry roughly twice the joint count. Highlanders presents the number as a decision rather than a limit: dexterity and reliability over strength, on a body that walks autonomously at 3 km/h, maps its surroundings with a 3D stereo camera and LiDAR, and brakes automatically on collision detection. Instructions are given in plain language and interpreted by a vision-language-action stack. The memorandum of understanding signed on 9 July 2026 is unusual in the industry: it is the first time a carmaker has agreed to mass-produce another company's humanoid. Mitsubishi contributes mass-production engineering, quality assurance, durability and safety design, and factory operations - exactly the disciplines a four-year-old startup cannot buy. The first robots go to work inside Mitsubishi's own plants, moving parts and assembling engines, before either company sells them elsewhere. Japanese press reports put planned capacity at up to 1,000 units a month; neither official release confirms that figure, or a price. The part that is easy to miss sits on Highlanders' own website. The company now describes its hardware as a data-collection loop for Kepler, its foundation model for physical intelligence - and Kepler is labelled "v1.0 - in development". Its target is a 10-billion-parameter world model trained on vision, force, tactile, teleoperation and simulation data, planning at about 5 Hz while a reflex layer runs at 500 Hz. The training cluster, named HISUI, is quoted at 640 NVIDIA B300 GPUs across 80 nodes and 24 petabytes of storage. Every one of those figures carries the word "expected". The company's stated data target - 26 million hours of real-world data - is for 2027 alone. So the timeline reads in two directions at once. The body is a known quantity, already through a beta programme that opened in September 2025 and an early access programme in Q4 2025, and it is now heading for a factory line with a 2027 date on it. The intelligence meant to run it is a specification, not a product. Mass production originally promised "within 2026" has become "early 2027", and the first customer will be the manufacturer's own assembly hall - which, for a robot that has to prove itself before anyone else buys it, may be the point.
Highlanders HL Human →On 10 October Alibaba's cloud stops serving its rivals' models. DeepSeek, Kimi, GLM and MiniMax all get the same replacement: Qwen
Model Studio, Alibaba Cloud's model platform, publishes its retirements as plain service notices, and six of them now converge on a single moment: 00:00 on 10 October 2026, Beijing time. After that hour the listed models stop answering; applications still calling them get nothing back. The list of Alibaba's own casualties is unremarkable housekeeping. The April notice retires qwen-turbo, qwen-vl-max, qwen-vl-plus, qwq-plus and qvq-max, all of them 2025-era products. The June notice goes further up the range and takes qwen3-max with it, along with qwen3-max-preview, qwen3.6-max-preview, qwen3-vl-flash and qwen3-coder-plus; buyers are pointed at qwen3.7-max at $2.5 and $7.5 per million tokens, at qwen3.6-flash, or at qwen3.7-plus. The interesting notice is the one about snapshots. Its table runs to roughly forty entries, and a large part of them are not Alibaba's models at all. DeepSeek-R1 and its 0528 revision, DeepSeek-V3, V3.1, V3.2 and the experimental V3.2-exp, the R1 distillations into Qwen 7B, 14B and 32B, Zhipu's glm-4.6 and glm-4.7, Moonshot-Kimi-K2-Instruct and kimi-k2-thinking, MiniMax-M2.1 — all of them are hosted third-party models that Alibaba sold access to, and all of them switch off on the same day. Against every one of those rows the recommended replacement column says the same thing: qwen3.7-plus, at $0.4 and $1.6 per million tokens below 256K of input, $1.2 and $4.8 above it. It is worth being precise about what this is and is not. None of these models is being discontinued by its maker. DeepSeek, Zhipu, Moonshot and MiniMax all publish open weights, and their models remain available from their own APIs and from other clouds; what ends on 10 October is Alibaba's willingness to host them. For customers who picked a Chinese rival's model precisely because it was available inside Alibaba's console, however, the practical effect is the same as a shutdown: migrate, move cloud, or accept Qwen. The direction of travel is easier to read from the replacement column than from the announcements themselves. A platform that spent 2025 advertising the breadth of its model catalogue is spending 2026 narrowing it to the models it makes. Speech is going the same way — a separate notice, also dated 10 October, closes Alibaba's hosted qwen3-asr-flash and qwen3-tts-flash endpoints and sends users to fun-asr and cosyvoice, which are not Qwen models either. Anyone running production traffic through Model Studio has under a month to check which of these names appears in their logs. The catalogue keeps profiles of the affected third-party models unchanged, because the models themselves are alive; the entry that ends is the one on Alibaba's price list.
DeepSeek-V3.2 →The company that invented the cobot is now defending it in court — across 17 countries at once
On 27 August 2026 Teradyne Robotics A/S filed a patent infringement case at the Unified Patent Court in Copenhagen against the German subsidiary of the Chinese manufacturer JAKA. The patents at issue are both software and hardware, and belong to the business unit that created this market: Universal Robots of Odense, whose UR5 established the collaborative robot as a product category in 2008. The venue is the point. The Unified Patent Court is a single court whose rulings take effect in 17 of the 18 EU member states that take part in it. A patent dispute that would once have been fought country by country can now settle the European availability of a whole product line in one judgment. Teradyne says the case covers a broad range of JAKA cobot models sold in the EU. It is the company's second intellectual property action in Europe this year. The earlier one ran in Germany against a subsidiary of Elite Robots and concerned copyright in Universal Robots' own software; the German court issued an injunction in Teradyne's favour. JAKA rejects the claim. The company says its technology comes from years of its own research, that it holds more than 300 granted patents worldwide, and that it commissioned two independent freedom-to-operate analyses before entering the European market. Why this matters beyond two company names: collaborative arms from Chinese manufacturers now undercut European ones sharply on price, and the incumbent's answer is being tested in a court whose reach is continental rather than national. No ruling has been issued, and nothing here is a finding against JAKA — the case has only been filed.
UR20 →Yaskawa promised five adaptive robots. The smallest one has no page to buy it from.
When Yaskawa announced the Motoman NEXT series in November 2023, it called it the industry's first industrial robot line able to judge a situation instead of replaying a taught path, and it named the lineup precisely: five models with payloads of 4, 7, 10, 20 and 35 kg, on sale from December 2023. The American launch in September 2024 repeated the same five and added two collaborative arms, NHC12 and NHC30. The catalogue looks different today. Yaskawa America's NEX series listing, checked on 6 September 2026, shows four robots: NEX7, NEX10, NEX20 and NEX35. There is no NEX4 tile, and the product address the series URL pattern would give it returns a 404 error page rather than a data sheet. The gap matters more than one missing tile. Adaptive behaviour — vision, force sensing, deciding what to do with a part that is not where it should be — is worth most on small, fiddly work: electronics assembly, laboratory handling, packing awkward little objects. That is the 4 kg class. A lineup that starts at 7 kg starts above the applications the technology was sold on. What the absence means is not stated anywhere. Yaskawa has published no withdrawal notice for the NEX4, and the model may still exist in other regions or as a special order; a product page can also vanish in a website rebuild. Until the company says otherwise, the honest reading is narrow: as of today, a North American buyer who wants the smallest adaptive Motoman has nowhere on Yaskawa's own site to start. Our catalogue profile of the NEX10, the middle size of the same family, records the same finding in its availability field.
Motoman NEXT NEX10 →Japan's companion robot gets 12% dearer — and the sticker was never the price
GROOVE X announced on 1 September 2026 that LOVOT 3.0, its wheeled companion robot, will cost more in two steps: the base colour goes from JPY 577,500 to JPY 599,500 on 26 October 2026, and to a planned JPY 649,000 in mid-January 2027. The company gives the reason plainly — rising prices of raw materials and parts, and higher manufacturing and logistics costs — and says the four-month, two-stage schedule is meant to give buyers time to decide. The older LOVOT 2.0 stays at JPY 449,900. That is a 12.4 percent rise on a robot that does nothing useful by design: it does not clean, carry or answer questions, it asks to be picked up. What the announcement does not say is that the purchase price has never been the whole cost. LOVOT cannot be used without a monthly plan, which carries the software, the cloud services and the servicing: JPY 9,900 a month at the cheapest tier, JPY 12,980 for the tier GROOVE X itself recommends, JPY 19,800 for full cover. Put the two together and the proportions change. At the recommended tier the subscription costs JPY 155,760 a year. Four years of it — roughly the interval at which the price list schedules a servo replacement pack — comes to JPY 623,040, more than the new purchase price of the robot itself. The announced increase raises the entry ticket by JPY 71,500; a single year of the plan costs more than twice that. None of this is hidden: the figures come from the company's own price list, which also states the servicing intervals and the colour surcharges of up to JPY 122,100. It is worth stating clearly all the same, because a headline price is what gets compared. Against a Chinese humanoid that keeps getting cheaper, LOVOT looks expensive. Against the four-year cost of keeping one alive, the sticker is the smaller half.
LOVOT 3.0 →Meituan flies 20 km with a winch at home and exports half the machine
Meituan runs the largest food-delivery drone fleet in the world, and it sells abroad under a separate brand: Keeta Drone, which operates in Dubai and Hong Kong. Its technology page lists exactly one aircraft, the Keeta Drone Gen 4: 7.2 kg empty, 2.4 kg of payload, 10 km of range, 10 m/s cruise. That is the standard model. In China the same fourth generation has a second member, the M-Drone 4L Winch, and the gap between them is not a detail. The winch version carries 4.5 kg instead of 2.4, flies up to 20 km fully loaded instead of 10, and lowers the parcel on a cable from a hover, so it does not need a prepared landing cabinet at the other end. Its stated weather envelope is heavy rain, moderate snow and Beaufort 6 wind, down to -20 °C. None of that is on offer to an export customer today. A partner signing up in the Gulf gets an aircraft with roughly half the payload and half the range of the strongest machine in the manufacturer's own fourth generation, and one that still depends on a fixed pick-up cabinet at the delivery end. There are unglamorous reasons why this can be so. A winch that pays out a cable over a street is a harder case for a regulator than a drone that lands inside a locked box on a rooftop, and export approvals are usually granted for one configuration at a time. Cold-weather figures also matter little in a country where the problem is 50 °C, not -20 °C. Still, the practical conclusion is worth stating plainly: the numbers quoted in coverage of Meituan's Chinese operations do not describe the aircraft on sale outside China. Anyone comparing drone delivery services should check which of the two is actually flying overhead.
Meituan M-Drone 4L Winch →Three generations of CogVideoX, three degrees of closing: Apache 2.0, a revocable licence, then no weights at all
Z.ai's video line has shipped three generations in two years, and each one has been harder to use freely than the one before. The catalogue now holds all three, so the sequence can be read in one place. On 6 August 2024 the producer published CogVideoX-2B, its first video model with downloadable weights: six seconds of 720x480 at eight frames per second, from 4 GB of video memory in the diffusers library and 3.6 GB with INT8 quantisation. It shipped under the producer's own restrictive document. Three weeks later, on 27 August 2024, one entry in the project's changelog did two things at once. It announced the larger CogVideoX-5B — same six seconds, same 720x480, about 5.6 billion parameters in the video transformer instead of 1.7, from 5 GB of memory in BF16 — and, in its closing line, moved the 2B to Apache 2.0. The bigger model did not get the same treatment. It went out under the CogVideoX Licence, which grants a copyright licence that is expressly revocable, permits academic use freely, and requires commercial users to register with the producer and stay below one million service visits a month. So the small model became genuinely free on the exact day the large one arrived without that freedom. That is the opposite of the usual pattern, where a producer opens the small model as a sample and holds the flagship back — here the flagship was released, but on terms the smaller model had just escaped. The third step removed the choice. CogVideoX-3, launched in July 2025, has no public weights at all. It is a hosted product: up to 3840x2160, five or ten seconds, 30 or 60 frames per second, an optional generated soundtrack, a flat $0.20 per video. Whoever built on the open versions cannot follow the line forward — they can only start again on someone else's weights, or pay per clip. The open two are not abandoned. In the 30 days to 4 September 2026, CogVideoX-2B was downloaded 20,387 times and CogVideoX-5B 17,161 times on Hugging Face — two years after release, and a year after the generation that replaced them stopped being downloadable at all.
CogVideoX-5B →Reasoning stopped being a product. It became a setting.
Two years ago, buying a model that thinks before it answers meant buying a different model. Today, at four of the five largest publishers, it means passing a parameter — and the separate reasoning lines are being switched off one by one. OpenAI's deprecation page now carries an end date for every surviving member of the o-series. o1, o1-pro, o3-mini and o4-mini shut down on 23 October 2026; o3 and o3-pro follow on 11 December. The recommended replacement in each case is a general model from the GPT-5.6 family, and for the two pro-tier models the replacement is written out as gpt-5.6-sol with reasoning mode set to pro. The premium reasoning tier is now literally a value in a request. Mistral got there first. Its entire Magistral line — the company's only reasoning products — was retired by 31 July 2026, and the migration targets listed by the manufacturer were not further reasoning models but ordinary models of the main line. DeepSeek reached the same place by a different route. Its current price list contains three models, all of them V4, and a single row headed thinking mode: both non-thinking and thinking modes are supported, with thinking as the default. The R1 line that made the company's name in January 2025 has no entry of its own. Alibaba's separate line ended earliest of all, and quietly. QwQ-32B-Preview appeared in November 2024 and QwQ-32B in March 2025; no third model ever followed. From Qwen3 onwards, thinking is a switch inside the general model. QwQ has a profile in this catalogue as of today — as the closing entry of a product category rather than the first entry of a family. Google is the exception that clarifies the rule, because it never made the separation in the first place. Gemini shipped thinking as a budget the caller sets, not as a model you choose, and so has nothing to wind down. For readers the practical consequence is narrow but real. Where a hosted reasoning model is switched off, its published benchmark results become unreproducible — the closed Magistral Medium 1.2 is already in that position. Where the weights were open, the model survives as a download after it stops being a product: QwQ-32B is still pulled tens of thousands of times a month, eighteen months after the line it belongs to quietly ended.
QwQ-32B →OpenAI is switching off Whisper's endpoint — and its own guide still sends subtitle work to it
On 26 August 2026 OpenAI put four transcription models on its deprecation list: whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize. All four leave the API on 26 February 2027. The recommended replacements are gpt-transcribe for recorded files and gpt-live-transcribe for live audio. The awkward part is in OpenAI's own documentation. Its transcription guide, checked on 3 September 2026, contains a table of specialised capabilities — and three rows of that table point at models that are now scheduled for shutdown. Speaker-labeled transcripts: use gpt-4o-transcribe-diarize. Word-level timestamps and srt or vtt subtitles: use whisper-1. Translation of a completed recording into English: use whisper-1. The two recommended successors are not listed for any of the three. So the platform is simultaneously recommending these models and removing them, without naming a like-for-like replacement for subtitle timing, speaker separation or speech translation. Anyone running a subtitling pipeline or a meeting-notes product on the OpenAI API has seventeen months to find out whether the successors cover their case, and OpenAI's documentation does not yet say that they do. There is one asymmetry worth spelling out, because it is the whole difference between an open model and a hosted one. Whisper's weights are published under the MIT licence as openai/whisper-large-v3. The shutdown does not touch them. Word timestamps, subtitle formats and English translation keep working on anyone's own hardware, for as long as they care to run them — the thing being switched off is OpenAI's convenience, not the model. The three GPT-4o transcription models have no such fallback: on 26 February 2027 they simply stop existing. Whisper is also still in the price list, at $0.006 per minute — the same rate as gpt-4o-transcribe and a third more than the model OpenAI now recommends. It has been the more expensive option for over a year, and it remains the one the documentation points to for the jobs the cheap model cannot do.
Whisper large-v3 →Mistral no longer sells a reasoning model: every Magistral now sits in the retirement table
In June 2025 Mistral AI launched Magistral, its own family of models that write out a chain of reasoning before answering. Fifteen months later the family is gone from the vendor's API. All six members — Magistral Small 1.0, 1.1 and 1.2, and Magistral Medium 1.0, 1.1 and 1.2 — appear in the retirement table of the lab's own documentation. The last two were switched off on 31 July 2026. The replacements named in that table are not another reasoning model. Users of Magistral Small are pointed to Mistral Small 4, users of Magistral Medium to Mistral Medium 3.5 — both ordinary mainline models. Reasoning did not disappear from the lab's line-up; it stopped being a product of its own. Since December 2025 the Ministral 3 family has shipped in three flavours of every size — base, instruct and reasoning — so a customer who wants deliberation picks a variant of the standard model instead of a different model. That is a pattern worth naming, because Mistral is not alone in it. A separate reasoning line made sense while the technique was new and expensive enough to be sold at a premium. Once it became a training stage that any model can be put through, keeping a parallel family meant maintaining two catalogues of the same sizes. What makes the case unusual is that switching a model off in the API did not switch it off in the world. Magistral Small 1.0, retired from the vendor's service on 30 November 2025, was downloaded roughly 80,700 times from Hugging Face in the thirty days to early September 2026 — more than six times the traffic of the final version 1.2, which is technically the better model. Devstral Small 2 makes the same point harder: withdrawn from the API on 31 March 2026, it pulled about 229,700 downloads in that same window, twelve times more than the 123-billion-parameter flagship it was released alongside. The counters cover thirty days only, and older releases have an advantage — tutorials, quantised forks and pipelines that were written once and still point at them. But the direction is clear enough. For a lab that publishes weights under Apache 2.0, retirement is a statement about what the company is willing to host, not about what people are running. That is why our profiles of Magistral Small 1.2 and Devstral Small 2, added today, are marked historical and at the same time describe how to run them — the service ended, the software did not.
Magistral Small 1.2 →Alibaba's video line has shipped four generations without publishing a single weight
Alibaba launched Wan3.0 on 24 August 2026: thirty seconds of video with sound in a single pass, and an input nobody else takes - a Word file, a spreadsheet, a slide deck, a PDF or a web link, which the model reads and turns into a film. It is a genuinely strong release. It is also the fourth Wan generation in a row with no weights published anywhere. That is the part worth stating plainly, because Wan built its name on the opposite. Wan2.1 in February 2025 and Wan2.2 in July 2025 went out under Apache 2.0 and became reference material for open video generation. Since then Alibaba's served line has moved through 2.5, 2.6, 2.7 and now 3.0. We checked the company's own Hugging Face account: the newest files there are refinements of Wan2.2 - Wan-Dancer-14B from 10 July 2026 and Wan2.2-Animate-2-14B from 14 July 2026. Searching the whole of Hugging Face for Wan2.7 returns nothing at all. The gap between what Alibaba runs for paying customers and what Alibaba gives away is now four generations wide. The numbers underneath make this less abstract. The most downloaded Wan repository is not a flagship: it is Wan2.1-T2V-1.3B, the smallest and oldest model of the line, at roughly 242 thousand downloads over thirty days, ahead of both 14-billion-parameter flagships. Second is Wan2.2-TI2V-5B, the version Alibaba built to run a five-second 720p clip in under nine minutes on a single RTX 4090. The audience for open video generation is visibly concentrated in the models that fit on a desktop - and the desktop-sized branch of the line is the one that stopped being updated first. We are not calling this a broken promise. Apache 2.0 grants no claim on future versions, and a company that has just raised about ten billion dollars in a Hong Kong share placement is entitled to decide what it sells and what it gives away. What we can say is what a reader choosing a video model today faces: the best Wan is rented by the second, the best Wan you can own is thirteen months old, and no announcement has stated that this will change. One caveat on our own reporting: Alibaba documents that Wan3.0 is billed per output second and that resolution changes the rate, but the price table on the company's site is rendered by script and we could not read it at source. The figures circulating in trade coverage - roughly five, ten and twenty cents per second for 480P, 720P and 1080P, with a 30 per cent discount to 23 September 2026 - are recorded in our profile as indicative rather than confirmed. The catalogue now carries three profiles from this line: Wan3.0, the closed flagship; Wan2.2-T2V-A14B, the last openly published flagship; and Wan2.2-TI2V-5B, the version for a single graphics card.
Wan3.0 →The humanoid boom shows up first in a sensor supplier's books
Robotics companies talk about production ramps; their suppliers have to ship the parts. RoboSense, one of the largest Chinese lidar makers, published its 2026 interim results on 26 August, and the robotics line of its business grew more than six times over. The company sold 282,600 lidar units into robotics in the first half of 2026, up 510.4 per cent year over year. Total shipments across all markets, cars included, were 719,200 units, up 169.6 per cent. Revenue reached roughly 1.02 billion yuan, up 30.2 per cent. Those three numbers do not move together, and the gap is the story. Shipments rose 170 per cent while revenue rose 30 per cent, which means the average lidar left the factory at a small fraction of last year's price. Volume is being bought with margin — the sensor that used to be the expensive part of a robot is becoming the cheap part. The robotics share is the sharper signal. Growth of 510 per cent from a supplier serving, by its own count, more than 3,400 robotics customers is not a pilot-project pattern; someone is building machines in quantity. RoboSense does not name those customers, does not break robotics down into humanoids, delivery robots and warehouse vehicles, and does not disclose gross margin in the release — so how much of this is humanoid legs and how much is logistics wheels remains unknown. What can be said is narrower and more useful than the usual forecast. Independently of anyone's shipment claims, a components maker has booked and reported the orders. The demand side of humanoid robotics is still mostly promises; the supply side has now filed numbers.
Renting out a humanoid in China paid for the machine in five days. Now it takes twenty - and the rent is no longer the point
China's robot rental market passed 1bn yuan (about $148m) in 2025 and iiMedia Research expects it to pass 10bn yuan this year - a tenfold jump. The price of the service collapsed just as fast in the other direction: a humanoid that cost 20,000 to 30,000 yuan a day to hire in early 2025 now goes for 2,000 to 5,000. The arithmetic behind that is unforgiving. In early 2025 a rentable machine cost around 100,000 yuan and five rental days paid for it. Today the same class of robot costs 30,000 to 50,000 yuan, but needs 10 to 20 rental days to break even. Operators who bought the pricier units, at 100,000 to 200,000 yuan, book only 10 to 15 rental days a month, so payback runs into months once maintenance and the operator's own wages are counted - figures reported by the Chinese business outlet 36Kr in May 2026. What collapsed was a particular kind of demand. The early money came from spectacle: robots hired to draw a crowd at a trade fair, perform at an event or stage a marriage proposal. That demand is episodic by nature. A machine that delivers room service in a hotel corridor or moves materials in a warehouse has a far stronger claim to recurring revenue, and the industry is now chasing the second kind. At this year's World Robot Conference in Beijing the shift was visible in the contract length. Shanghai-based Futuring Robot rents a household robot for 3,000 yuan a month - a machine that costs the company more than 20,000 yuan to build. It tried one-day trials and weekly rentals of about 700 yuan before settling on the monthly model. Bookings already run into December, and customers are deliberately capped at one month each so the machines circulate through more homes. That cap is the tell. Co-founder Louis Shen told China Economic Net that low-priced household rental schemes are not primarily designed to make money from rent at all: their strategic value is the chance to collect data from real homes and use it to train world models. It inverts the normal economics of leasing. A conventional lessor wants an asset to stay with one customer as long as possible; a robot company may want the opposite, because a home that has already been mapped teaches it nothing new. The rental contract doubles as a field test in the one environment embodied AI cannot rehearse - furniture moves, objects land in the wrong places, and children and pets do not follow a demo script. The state is pushing in the same direction. In June the Ministry of Industry and Information Technology and SASAC called for real-world testing of humanoids and embodied-intelligence systems, measured on task success rates, efficiency, safety and economic feasibility, and explicitly encouraged Robot-as-a-Service models based on usage-based payments and operating leases. The stated aim is more than 100 high-value application scenarios and deployment capacity in the tens of thousands of units by the end of 2026. Beijing's Economic-Technological Development Area subsidises 10 per cent of rental costs for qualifying projects, capped at 3m yuan per company a year, and backs insurance for humanoid robots. Some context on scale: China produced 635,056 industrial robots in the first seven months of 2026, up 28.5 per cent year on year, and 12.14m service robots, up 12.2 per cent. Against that, a 10bn-yuan rental market is small. It is worth watching anyway, because it is where the industry is testing a different answer to the adoption problem - one that treats the obstacle not as the price of the machine but as the customer's unwillingness to bet capital on hardware whose capabilities are still changing every few months.
Richtech withdrew two years of accounts - and quietly two robot lines with them
Richtech Robotics, the Las Vegas maker of service robots listed on Nasdaq as RR, told investors on 7 August that its financial statements for the fiscal years ended September 2025 and 2024, plus four interim periods, should no longer be relied upon. The audit committee reached that conclusion on 9 June while reviewing the March 2026 quarter. The restated annual report was filed the same day as an amended quarterly report and a late one; a formal notice of late filing followed on 14 August, and the delayed quarterly arrived on 19 August. The errors are accounting, not operational. The company says it mishandled warrants issued in equity financings - placement agent warrants that should have counted as non-employee share-based pay, and pre-funded, common and inducement warrants that should have sat on the balance sheet as liabilities and been remeasured at fair value each period. It also failed to treat a standby equity purchase agreement signed with YA II PN in February 2024 as a derivative, misclassified cost of revenue and research spending as general and administrative expense, booked inventory purchases as property and equipment, and shortened the useful lives of certain intangibles. For readers of a robot catalogue, the more interesting sentence sits two pages earlier, in the part of the filing that describes what the company actually sells. "The Skylark line of robots has been discontinued to refocus resources on products with better long-term profitability," it reads. "The Medbot line of robots has been discontinued while we develop new robotic technologies to better service the healthcare sector." That is two whole categories gone. Skylark was a modular hotel robot: one self-driving chassis that rode the building's lifts, with a delivery module for room service and a cleaning module for corridors, plus linen, waste, security and disinfection modules that were advertised and never shipped. Medbot was the hospital courier - four lockable compartments opened by PIN or fingerprint, a machine the company said moved eight to nine thousand deliveries a month in a five-robot fleet. Neither was ever given a published price. What remains is six lines in two pillars: Matradee, ADAM and Scorpion in restaurants and bars; Titan, DUST-E and the newly introduced Dex humanoid in industry. Hospitality delivery now falls to the Matradee restaurant line; healthcare has no successor product at all. The company is not behaving like a business in retreat. On 21 August its board authorised the repurchase of up to $12 million of Class B shares over the following year, announced four days later. The same compensation committee meeting granted the chief executive, chief financial officer and chief operating officer a healthcare stipend of $750 per fortnightly paycheck each, about $19,500 a year apiece. wujec.ai has updated both profiles: Medbot is now marked as withdrawn, and Skylark - which the catalogue did not have - has been added as a retired machine rather than left out. A robot that was sold and then dropped is part of the record.
Richtech Skylark →Ant's trillion-parameter flagship has 7,000 downloads. Its smallest model has a million
InclusionAI, the open-source arm of Ant Group, publishes at both ends of the scale: trillion-parameter reasoning models and 16-billion-parameter models meant for a single accelerator. The download counters on Hugging Face show which end the world actually uses, and the gap is not close. As of 29 August 2026, LLaDA2.0-mini has been downloaded 1,013,751 times since its weights appeared on 25 November 2025. Ring-2.6-1T, the company's trillion-parameter reasoning flagship released on 14 May 2026, has 7,237. That is a ratio of roughly 140 to 1 in favour of the model that is about sixty times smaller. Size alone does not explain it, because the same pattern shows up inside each line — and in both cases it points backwards, to the older release. Ring-2.5-1T, the previous trillion-class model, has 175,004 downloads across 199 days, an average of 879 a day; its successor Ring-2.6-1T averages 68. Over the last 30 days the older model runs at 312 downloads a day and the newer one at 16, so the gap is widening rather than closing. In the diffusion line, LLaDA2.0-mini averages 3,673 downloads a day and its February successor LLaDA2.1-mini averages 1,429. The usual excuse — that a newer model has had less time to accumulate downloads — does not apply here, because these are daily rates measured from each model's own publication date. Nor is it a packaging problem: both trillion-class models ship natively in FP8 and contain the same number of FP8 parameters, 1,009,494,523,904. Neither is resold by a single provider on OpenRouter. An earlier suspicion, that only the older one was documented in the official SGLang deployment cookbook, turns out to be wrong: the cookbook lists Ring-2.5-1T and Ring-2.6-1T alike. One caveat on the counters themselves. Hugging Face reports both a 30-day figure and a lifetime figure, and only the lifetime figure allows a fair comparison across models with different release dates; the widely quoted 30-day number makes any recent release look weak by construction. Downloads also measure trial and tooling traffic, not production use — a file pulled by a benchmarking script counts the same as one pulled by a company. What the numbers do suggest is that for open weights, the constraint is hardware, not capability. A model that fits on one accelerator gets tried; a model that needs a rack gets read about. The catalogue has today added a profile of LLaDA2.0-mini, the most downloaded model this publisher has ever released.
LLaDA2.0-mini →The same hour of audio costs 10 cents or 61 — and no benchmark covers all the sellers
Three transcription vendors entered this catalogue today, and putting their price lists side by side produces a number that is hard to explain: an hour of recorded audio costs 10 cents at Soniox and 61 cents at Gladia. Between them sit Speechmatics at 12.9 cents for its newest model and AssemblyAI at 21 cents. Same task, same month, a six-fold spread. Within a single vendor the direction is not consistent either. Speechmatics released Melia 1 on 17 June as its broadest model — it takes no language setting at all and switches between 55-plus languages inside one recording — and priced it at roughly half of Standard, the model most of its customers run, and a third of Enhanced, its most accurate one. AssemblyAI moved the other way: its June flagship costs 21 cents against 15 for the older model it sits above, which also covers far more languages. What is missing is any way for a buyer to check whether the expensive hour is a better hour. Every set of comparative numbers published this summer was produced by one of the sellers, running its rivals' systems itself. Gladia's June campaign for Solaria-3 measured AssemblyAI, ElevenLabs, Deepgram, Mistral and Speechmatics — and excluded Soniox, the cheapest of them, for lack of data. Speechmatics reports beating Deepgram and Microsoft on 91 percent of FLEURS languages, on its own runs. Soniox publishes no error rate for its current model at all, only a list of things it does better than the model it replaced. One vendor did something the others did not. Gladia printed the benchmarks on which its new model is worse than its old one — 32 percent worse on recordings of parliamentary debate, 36 percent worse on read audiobooks — and told customers with that kind of audio to stay on the older model. It is a narrow admission, and it is still the only published number this month that works against the company that published it. For readers choosing a supplier, the practical consequence is that the price list is the only figure in this market that has been independently fixed. Everything about accuracy is currently a claim made by an interested party, on test audio of its own choosing.
Solaria-3 →The only humanoid you can actually order is being bought at a lower price than it asked for
SoftBank is in talks to take a majority stake in 1X Technologies, the Norwegian maker of the home humanoid NEO, in a deal that would value the company at about 6 billion dollars. The talks were first reported by The Information on 26 August 2026 and relayed by Reuters; SoftBank declined to comment and 1X did not respond to a request for comment. Terms may still change. The number is the story. According to reporting on last year's fundraising, 1X was seeking a valuation of roughly 10 billion dollars for a round of about 1 billion. A controlling stake at 6 billion is therefore not a step up in the middle of a humanoid boom — it is a markdown of about 40 per cent against the company's own asking price, and it comes with a loss of control. Set that against what the money is buying. NEO is one of very few humanoids a private customer can order at a published price, and its 2026 hand is the most detailed piece of hardware 1X has documented: 25 degrees of freedom, force control on every axis, fingertip sensing that measures shear so a slipping object can be re-gripped. Three days before this report, XPeng's robot unit — which has no customers at all — was valued at 6.3 billion dollars in a funding round. A shipping consumer product is being priced below a pre-revenue subsidiary. For SoftBank the logic is continuity rather than novelty. The group agreed in 2025 to buy ABB's robotics division for 5.4 billion dollars, and it has been assembling automation assets around its AI investments; 1X would add a consumer-facing humanoid and a home dataset to an industrial portfolio. OpenAI's Startup Fund led an early round in 1X in 2023 and OpenAI itself is reported to have looked at buying the company in 2025 before those talks ended. Nothing has been signed. But if the price holds, the first serious valuation event for a humanoid that ordinary buyers can order will have set the market lower, not higher.
NEO →Agility goes public on $300m of orders and 65,000 robot hours — the equivalent of 32 human working years
Agility Robotics announced on 27 August that its chief financial officer will meet institutional investors at three conferences in September. It is a routine announcement, but it is the first public step of something that is not routine at all: the American maker of the Digit humanoid is heading for the stock market not through an initial public offering, but through a merger with Churchill Capital Corp XI, a listed shell company with no business of its own. The terms were set out in June. The transaction values Agility at a pre-money equity value of $2.5bn and is expected to hand the company more than $620m in gross proceeds: $420m held in the shell's trust account, assuming no shareholder asks for their money back, plus roughly $200m of fresh institutional money committed at $10 per share. The combined company is to trade under the ticker AGLT. What stands behind that valuation is, above all, an order book. Agility says it has secured more than $300m of multi-year orders for Digit v5, its next-generation robot — but the company itself adds that those orders are subject to the realisation of certain contractual milestones. In other words, they are commitments, not revenue. Agility has not published a revenue figure at all. Digit is deployed today with Schaeffler, GXO, Toyota Motor Manufacturing Canada and Mercado Libre, and the company names a pipeline of more than 30 further customers. One hard operating number does appear in the documents, and it is worth reading slowly. Across deployment commitments at nine customer facilities, Digit has accumulated more than 65,000 hours of operation — the total for the whole fleet, over years of work. A full-time human year is about 2,000 hours. The entire commercial history of the robot therefore amounts to roughly 32 human working years. Agility's RoboFab plant is designed for up to 10,000 robots a year. This is the second humanoid maker to approach the public markets this month. Unitree debuted in Shanghai on 19 August and closed its first session valued at 201 times its previous year's sales — but Unitree sells robots by the thousand and reports revenue. Agility's case rests on a different argument: that Digit v5 will be the first humanoid safe enough to work beside people rather than behind a fence, and that this is what unlocks the market. One caveat belongs in the reader's mind for the months ahead. In a merger of this kind, the shell company's shareholders may demand their money back before the deal closes, and the $420m in the trust account can shrink sharply. The $620m figure is a ceiling, not a certainty.
Digit v5 →The company that put humanoids in car factories now sells one that holds your hand — 13,361 ordered
UBTech Robotics spent two years building its reputation on humanoids that carry parts around BYD and NIO factories. This summer it quietly split into three companies' worth of product lines, and the newest one is not industrial at all. On 22 June the manufacturer unveiled Walker C1, a service humanoid for lobbies and exhibition halls: 164 cm, 55 kg, 53 degrees of freedom, with a 3D-printed lattice used as the muscle framework and a quoted top speed of at least 3 m/s. It debuted at the fourth China International Supply Chain Expo in Beijing as the event's first official humanoid representative, working the floor as a guide and receptionist. Eight days later, on 30 June in Shenzhen, came something else entirely. The UWORLD U1 series is a companion robot with biomimetic silicone skin, 88 degrees of freedom, a proprietary dual-pivot cervical spine, and a language model built to recognise more than 20 fine-grained emotional states — UBTech claims above 90 percent accuracy, a 500-millisecond response and lip-sync held within 20 milliseconds. It comes in three forms: U1 Lite as a semi-torso, U1 Pro as a full body, and U1 Ultra as the high-dynamic version. Prices start at 119,800 RMB, roughly 16,500 US dollars. The number that stands out is not a specification. UBTech reported 13,361 cumulative orders for the U1 series at the launch event — for a machine whose purpose is company, not work. UBTech also pledged to donate 100 customised units during 2026 to isolated seniors, children separated from their parents and families in hardship. The software claims are where a reader should stay careful. UBTech describes a fast-and-slow brain architecture running models in the hundreds of billions of parameters, an Agent Memory OS that carries memory between conversations, and a care engine that acts on its own initiative rather than waiting to be addressed. Privacy is presented as three layers: local-first processing, minimal cloud dependency and hardware cut-offs under the user's control. That last point is not decoration — this is a machine with cameras and microphones pointed into a living room. What UBTech has not published is as telling as what it has. For neither Walker C1 nor any U1 variant has the company disclosed runtime, battery capacity or compute platform; for the U1 series it publishes no heights or weights at all, only series-level figures; and it has not said how many of those 13,361 orders have actually been delivered. Both machines are now in the wujec.ai catalogue with those gaps marked as gaps.
UWORLD U1 →XPeng's robot unit is valued at $6.3bn before it has a single customer
XPeng announced on 24 August 2026 that its humanoid robotics business had raised more than US$900 million at a post-money valuation of over US$6.3 billion. By the company's own timeline, that valuation was set roughly a year and a half before the first robot is due to reach a paying customer. The round was led by IDG Capital, with Gaorong Ventures participating and Tencent and Alibaba coming in as strategic investors. XPeng calls it the largest single private financing round ever recorded in China's embodied AI industry. The carmaker keeps control: the robotics business stays a consolidated subsidiary. The schedule attached to the money is the part worth reading twice. Mass production of the IRON humanoid is expected at the end of 2026. The first units go to XPeng's own stores and campuses. Official launch and customer deliveries, in China and abroad, are scheduled for 2027. That sequence means the $6.3 billion is not a multiple of anything. When Unitree listed in Shanghai five days earlier, its closing price worked out to 201 times last year's sales — an extreme number, but a number, because there were sales to divide by. Here there is no denominator. The first party to take delivery of an IRON is XPeng itself. What the money buys is a machine the company describes as 76 degrees of freedom across the body and 21 in each hand, wrapped in a proprietary fully enclosed flexible lattice structure and driven by three in-house Turing AI chips rated at a combined 2,250 TOPS. Those chips are the strongest argument for the valuation: they are the same silicon XPeng already builds for its cars, which makes the robot programme an extension of an existing supply chain rather than a start from zero. One caveat on the specifications. XPeng's Chinese product page for IRON still states 22 degrees of freedom per hand, while the English-language release issued this week says 21. Both are first-party documents published in the same week. We keep the figure from the product page in our profile and will update it when the company reconciles the two. No price for IRON has been published, no order book has been disclosed, and no independent party has tested the robot.
XPeng IRON →o3 disappears from ChatGPT on Wednesday — and keeps selling in the API for another 107 days
On 26 August 2026 OpenAI o3 stops being selectable in ChatGPT. The date was set on 28 May, when OpenAI gave the model a 90-day sunset alongside GPT-4.5, which went on 27 June. Both were paid-tier-only options hidden behind model settings by then. The same announcement contains the sentence that most coverage drops: "These changes apply to ChatGPT only; there are no changes to the API." In the API, o3 is scheduled to shut down on 11 December 2026, together with o3-pro and the whole first generation of GPT-5 — gpt-5, gpt-5-mini, gpt-5-nano and gpt-5-pro. Every one of them is pointed at gpt-5.6-sol. That is 107 days between the two deaths of the same model. The gap can be much wider. GPT-4o was removed from ChatGPT on 13 February 2026 and is still on sale in the API today at $2.50 per million input tokens, with no shutdown date announced at all. Only its oldest snapshot, gpt-4o-2024-05-13, is scheduled to go, on 23 October — and the model's default snapshot has been gpt-4o-2024-08-06 for two years. A model can be gone from the product the public knows and still be a supported production endpoint six months later. That is why the autumn calendar reads differently depending on which side you are on. For consumers, the visible loss is o3 on Wednesday. For anyone shipping software, the dates that matter are 23 October — when gpt-3.5-turbo, gpt-4, gpt-4 Turbo, o1, o1-pro, o3-mini, o4-mini and gpt-image-1 go for good — then 1 December for the older image models, 11 December for o3 and the GPT-5 family, and 20 January 2027 for the legacy audio and realtime line. OpenAI now publishes the notice periods behind those dates: at least six months for generally available models, at least three for specialised variants such as chat or Codex builds, and as little as two weeks for anything with "preview" in its name. It is a commitment about warning time, not about lifespan — and the ChatGPT calendar it does not cover has been the faster of the two all year.
OpenAI o3 →Honor says five human world records fell to its robot — and that the point of the robot is the phone
Honor is a smartphone company. It has no robotics product line, no robot for sale and no price list for one. Yet at the second World Humanoid Robot Games in Beijing it fielded two machines, and on 22 August 2026 it put up a campaign page of its own to tell the world how they did. The page claims five human world records beaten by the larger robot, called Lightning: 9.32 s over 100 m, 39.45 s over 400 m, 2 minutes 30 seconds over 1500 m, 50 minutes 26 seconds over the half marathon, and a peak speed of 14.5 m/s. The page is also where the machine finally acquires a manufacturer designation, something no event report has used: the plaque carried at the athletes' oath ceremony reads HONOR Robotics D1. Those five figures need separating. The 400 m mark is a competition result — Honor states it was set in the large-robot final at the games. The 100 m figure is not: in the 100 m final at the same games Lightning was timed at 9.47 s and finished second, beaten on the line by Tiangong Ultra at 9.39 s. The 9.32 s comes from Honor's own runs. A record card and a results table are different documents, and this one is the record card. The sentence that explains the whole programme sits lower down the page, under the smaller of the two robots. Yuanqizai took silver in the 400 m small-robot final at 48.47 s, and Honor adds: the technical strength tested on the world stage will return, in full, in the Honor Magic9 series. Not in a robot. In a phone. That closes a loop this catalogue has already documented from the other end. Lightning's most distinctive component is its cooling: a liquid loop with a suspension pump turning above 20,000 rpm and moving up to 6 litres a minute, which is handset thermal engineering scaled up so the actuators can hold full power through a long, hot run. Honor built the robot partly to find out how far that technology bends. Now, by its own account, the answer goes back into the handset. It is a more honest statement of purpose than most humanoid programmes offer, and it is worth taking at face value rather than as modesty. A machine that runs 400 m and nothing else is not a product and was never meant to be one. It is a test rig with a racing number on it — and the thing being tested is not the robot.
Honor Lightning →Unitree closed its first trading day worth 201 times last year's sales — its own filing puts growth at 36%, not 333%
On 19 August 2026 Unitree Robotics became the first humanoid robot maker to trade on a mainland Chinese exchange. It priced its offering on Shanghai's STAR Market at 150.80 yuan a share, sold 40.45 million shares — a tenth of its enlarged capital — for about 6.1 billion yuan, and entered the session valued at roughly 61 billion yuan. The market disagreed with that price immediately. The stock opened at 1,100 yuan, 629% above the offer, briefly worth about 445 billion yuan (US$66 billion). It closed at 845 yuan, still up 460%, at about 342 billion yuan. Around 23.2 billion yuan of shares changed hands in a single day, while the STAR Market Composite fell 7.2% and the Shanghai Composite 2.4%. What the buyers were paying for is set out in the prospectus. Revenue reached 1.70 billion yuan in 2025, against 392.77 million yuan the year before — a 333% jump. Net profit was 278.21 million yuan. At the closing price, the company was therefore worth roughly 201 times its annual sales and about 1,230 times its annual profit. The same filings describe a business that is already slowing. Unitree guides for first-half 2026 revenue of about 1.1 billion yuan, growth of 35.6% to 45.4% year on year — roughly a tenth of the rate that produced the debut. First-quarter profit excluding one-off items fell 52.6% as spending on research and marketing rose. 2025 was also the year the robot dog stopped being the main product. Humanoid sales of 867.8 million yuan overtook the quadrupeds, which slipped to about 42% of revenue. Unitree shipped more than 5,500 humanoids, more than any other maker in the world, on top of more than 33,000 quadrupeds sold to date. Divide the humanoid line by the units and the average machine brought in about 158,000 yuan — roughly the price of one G1 with options, not of an industrial platform. The risk the prospectus names is American. Overseas customers provided 43.65% of main-business revenue, 731.66 million yuan, and the United States alone 13.3% — in the same season that US regulators moved to bar new equipment approvals for foreign-built humanoid and quadruped robots. Unitree says the products it already sells there keep their authorisation. The sharpest number of the day was not the share price. Meituan, the food delivery company whose own delivery drones we described yesterday, held 8.7% of Unitree after the listing: close to 30 billion yuan at the closing bell, about 70 times what it originally paid. Founder Wang Xingxing, 36, held 121.4 million shares worth roughly 103 billion yuan.
Unitree G1 →OpenAI is switching off its entire image line by 1 December — and the cheapest model dies with it
OpenAI's deprecation page now schedules the end of every image model the company sells except one. gpt-image-1 goes dark on 23 October 2026; gpt-image-1.5, gpt-image-1-mini and chatgpt-image-latest all follow on 1 December. Every one of them names the same replacement: gpt-image-2. For most users that is a price cut. Priced per million image tokens, the 2025 original gpt-image-1 is the dearest model in the whole line at $10 in and $40 out. gpt-image-1.5 charges $8 and $32. The survivor, gpt-image-2, charges $8 and $30 — 25% less on output than the model it replaces, which is the opposite of what the industry has been doing this year. The exception is the one model built for people watching costs. gpt-image-1-mini charges $2.50 per million image tokens in and $8.00 out. It has no successor of its own: the migration table sends it to gpt-image-2 as well. That is 3.2 times more on input and 3.75 times more on output, for anyone who chose the mini precisely because it was cheap. The pattern is becoming familiar. Cheap tiers are announced as an entry point, then quietly folded into the flagship when the line is consolidated — and the bill for the migration lands on the users who were most price-sensitive to begin with. Batch pricing halves all of these figures, but it halves them for both the old model and the new one, so the ratio does not move.
Two humanoid shipment reports, one day, one impossible number: China alone ships more than the whole world
Two reports on the same industry landed on 20 August 2026, hours apart, and they do not fit together. At the World Robot Conference in Beijing, the China Humanoid Robotics and Embodied Intelligence Committee of 100 released its 2026 Humanoid Robot Industry Development Report. Its headline figure: China shipped more than 40,000 humanoid robots in the first half of 2026, lifting the country's share of global shipments to 97 percent. Taken literally, that implies a world market of roughly 41,000 machines. The same day, Counterpoint Research published its own first-half tally: global humanoid shipments "topped 22,000 units", up nearly 300 percent year on year. Counterpoint breaks the figure down by vendor — AGIBOT first with about 9,700 units and over 43 percent share, Unitree second with more than 7,000 units and 31 percent, followed by Galbot, UBTECH and Leju, with the top five accounting for 86 percent of shipments. So one report says China alone shipped 40,000. The other says the entire planet shipped 22,000. The Chinese figure is not merely higher than the independent analyst's China figure — it is nearly double the analyst's number for the whole world, China included. Neither document explains the gap, and the committee's report is not published in full, so its counting method cannot be inspected. Both are quoted in the trade press without the contradiction being noted. There is a second detail worth more than the headline. Counterpoint's application breakdown shows that entertainment and performance robots, together with machines bought to generate training data for research, still account for more than 60 percent of all humanoid shipments. Intelligent manufacturing takes 13 percent and warehousing and logistics 5 percent. The Chinese report describes an industry that has "moved past small-batch trials into routine deployment at scale"; the shipment mix suggests that most of what ships is still rented out for shows or wired up to produce data — not put to work. Counterpoint expects global shipments to pass 50,000 units for the full year 2026. Editorial note: we quote the 40,000 figure as reported, because the underlying report has not been made public. Where the two sources conflict, we give both rather than choosing between them.
Every xAI model that replaced another in 2026 costs more — text, image, video and voice alike
In a market that talks constantly about the falling cost of intelligence, xAI spent 2026 moving in the opposite direction — and it did so in every product line it sells. Text. Grok 4.3, released on 30 April 2026, is priced at $1.25 per million input tokens and $2.50 per million output tokens. Grok 4.5 arrived on 8 July at $2.00 and $6.00 — 60% more for input, 140% more for output. Grok 4.6, released on 12 August, kept those numbers unchanged, so the flagship rate has not come back down. Both newer models double their rate again above a 200,000-token prompt, and the higher rate applies to the whole request, not just the part above the threshold. Images. Grok Imagine Image, from 28 January, generates a picture for $0.020 at both 1K and 2K. Grok Imagine Image 2.0, from 7 August, charges $0.04 for the same job — twice as much. The separate quality tier introduced on 3 April sits at $0.05. Video. The original Grok Imagine Video, also from 28 January, costs $0.050 per second of output at 480p. Version 1.5, released on 31 July, costs $0.080 per second — 60% more, and, as we reported earlier this week, it is the one model xAI's own batch discount refuses to cover. Voice. Grok Voice Think Fast 1, from 23 April, is $0.05 per minute of audio. Think Fast 2, from 29 July, is $0.08 per minute — again 60% more, on top of the same $0.004 text input charge. The cheapest text model xAI sells is not the newest one either: Grok Build 0.1, from 29 May, runs at $1.00 and $2.00 per million tokens, below every general-purpose Grok on the list. None of this makes xAI unusual on quality; it makes the industry's assumed direction of travel worth checking. Anthropic has held the Opus rate steady across five releases since November. Mistral's own mid-tier moved the other way in dramatic fashion — Medium 3.5 costs 3.75 times what Medium 3 costs — but its larger Mistral Large 3 is cheaper still, at $0.50 and $1.50. Reading a price list by version number is the fastest way to overpay.
Grok 4.6 →Mistral retires Medium 3 in eight days — the replacement it names costs 3.75 times more per token
On 31 August 2026 Mistral AI switches off two models at once: Mistral Medium 3 (API name mistral-medium-2505, released May 2025) and Mistral Medium 3.1 (mistral-medium-2508, August 2025). Both were marked deprecated on 22 May 2026, and both are billed at $0.40 per million input tokens and $2.00 per million output. The company's documentation names a single replacement for them: Mistral Medium 3.5. That replacement is listed at $1.50 input and $7.50 output. The multiplier is the same in both directions — 3.75 times. A team that keeps its prompts, its volumes and its code exactly as they are, and does nothing but follow the migration path printed in Mistral's own table, will see its bill for this model line grow by 275% on the first of September. The part that is easy to miss is that the cheaper option is upwards, not sideways. Mistral Large 3, the company's larger and older model from December 2025, costs $0.50 input and $1.50 output — a third of Medium 3.5's input price and a fifth of its output price. On Mistral's public price list the ordering is inverted: the mid-tier model is the expensive one, and the model above it in the naming scheme is cheaper than both the model it sits above and the model it replaces. Below them, Mistral Small 4 runs at $0.15 and $0.60. The inversion is worth noting because of how Medium 3 was sold. Mistral launched it in May 2025 under the headline "Medium is the new large", with a pitch built almost entirely on price: roughly eight times cheaper than comparable models, deployable on four GPUs, at or above 90% of Claude Sonnet 3.7 on the company's own benchmarks. Fifteen months later the tier that was created to be the cheap one is the one that costs the most per token in the range. None of this makes Medium 3.5 a bad model — it is newer, and Mistral positions it as the stronger of the two. But the migration note in the documentation says only "use Mistral Medium 3.5 for new integrations", and says nothing about what that costs. Anyone still calling mistral-medium-2505 or mistral-medium-2508 has eight days to decide whether the named successor, the larger model or the smaller one is the right destination.
Mistral Medium 3 →Five Opus releases, one price: Anthropic's flagship has not changed its rate since November — and the dearest Claude on sale is not a flagship
Anthropic has released five Opus-class models since 24 November 2025: Opus 4.5, 4.6, 4.7, 4.8 and Opus 5. Every one of them is billed at exactly five dollars per million input tokens and twenty-five per million output. Eight months, five model numbers, one unchanged rate. That is the second stable plateau in the tier's history, and the first one lasted even longer. Claude 3 Opus launched in March 2024 at $15 / $75. Opus 4 in May 2025 charged $15 / $75. Opus 4.1 in August 2025 charged $15 / $75. Twenty months and three flagships went by without the price list moving. It moved once, on 24 November 2025, when Opus 4.5 cut the rate by two thirds — and it has not moved since. The practical consequence is easy to state: the current flagship, Opus 5 from 24 July 2026, costs one third of what the first Opus cost in March 2024, and there was no gradual glide between the two numbers. There was one step. The middle tier tells the same story with different digits. Claude 3.5 Sonnet, 3.7 Sonnet, Sonnet 4, Sonnet 4.5 and Sonnet 4.6 were all sold at $3 / $15 — five models across twenty months at a single rate. Sonnet 5 broke it on 30 June 2026 at $2 / $10, and Anthropic then cancelled the increase back to $3 / $15 that it had scheduled for 1 September. Only the small tier has ever gone the other way: Claude 3 Haiku cost $0.25 / $1.25 in March 2024, and its successor 3.5 Haiku arrived at $1 / $5 — a fourfold rise that Haiku 4.5 has kept. The part that is easiest to miss sits at the top of the list. The most expensive Claude a customer can buy today is not Opus 5. Claude Fable 5 and Claude Mythos 5, all published on 9 June 2026, are billed at $10 / $50 — double the flagship. Opus 4.8 hints at what that premium buys: its own fast mode is priced at $10 / $50 for roughly 2.5 times the speed. On Anthropic's price list, in other words, the second five-dollar step is not sold as a better model. It is sold as a faster one. The figures in this piece are taken from the release notes and price list entries recorded in the wujec.ai catalogue profile of every Claude model, and can be checked profile by profile.
Claude Opus 5 →MiniMax's most downloaded model is the one you may not sell anything with
MiniMax M2.7 is pulled from Hugging Face roughly 909,000 times a month — more than any other text model the Chinese lab has published, and more than four times the traffic of its own successor M3. It is also the only one in the family that forbids commercial use. The licence file shipped with the weights is titled "Non-Commercial License". Personal use, self-hosted deployment, research and experimentation are expressly free of charge. Everything else is not: selling a product or service that relies on the model, putting it behind a paid API, or deploying a fine-tuned derivative for commercial purposes all require prior written authorisation from MiniMax, requested by e-mail. Anyone who obtains that permission must also display "Built with MiniMax M2.7" on the product. That term is an outlier rather than a direction of travel. MiniMax M2 (October 2025) and M2.1 (December 2025) were released under MIT with a single added clause: a commercial product built on the model must show the model name in its interface. In M2 that duty starts only above 100 million monthly active users or $30 million in annual recurring revenue; in M2.1 it applies with no threshold at all. M2.5 (February 2026) moved to the company's own MiniMax Model License, which still allows redistribution but requires a fixed attribution notice to accompany every copy. M2.7, released a month later, closed commercial use altogether. M3, in June 2026, reopened it under a community licence that permits commercial deployment in exchange for visible attribution. The practical consequence is easy to miss, because nothing about the download page signals it. Four models in this line share the same body — 62 layers, 256 experts, eight routed per token, and 228,689,764,864 parameters in M2, M2.1 and M2.7, to the byte. A team that swaps one for another as a drop-in upgrade changes its legal position without changing a line of code. The strongest freely reusable model in the line is M2.5; the strongest one overall, by the producer's own benchmark table, is the one that needs a signature. wujec.ai has today added catalogue profiles for M2.1, M2.5 and M2.7, each stating the licence terms in full.
MiniMax M2.7 →Mistral built its name on freely licensed weights — its biggest one is still the model it published in April 2024
Mistral AI became a recognised name in September 2023 by publishing a 7.3-billion-parameter model under Apache 2.0 — the licence that lets anyone download, fine-tune and commercially resell the result without asking. Three months later, Mixtral 8x7B repeated the trick at a size that mattered: 46.7 billion parameters stored, 12.9 billion spent per token, benchmark scores at the level of Llama 2 70B and of the GPT-3.5 base model then serving free ChatGPT. In April 2024, Mixtral 8x22B scaled the same sparse design to 141 billion parameters, 39 billion of them active per token, with a 64,000-token window and native function calling. It, too, went out under Apache 2.0. That April 2024 release is still the largest model Mistral has ever opened. Everything above it changed terms: Mistral Large 2.1, from November 2024, shipped its weights under the Mistral Research License with a separate commercial licence sold on top, and the flagships that followed were never opened the same way. The company that made permissive licensing its signature has kept publishing open weights — but at the small and medium end, while the top of its range moved behind commercial terms. The practical consequence is that Apache 2.0 cannot be revoked. Mixtral 8x22B has disappeared from Mistral's own API line-up, which today lists Large 3, Medium 3.5, Small 4 and Ministral 3, and the company has never announced a shutdown date for it. That is not a problem for anyone relying on it: the weights sit on Hugging Face, third-party providers still serve them, and a licence granted under Apache 2.0 stays granted regardless of what the vendor's catalogue says. wujec.ai has now catalogued all three of these open releases.
Mixtral 8x22B →Cohere's oldest model costs six times more than its newer replacement — and has no shutdown date to force anyone off it
Cohere's pricing page carries a quiet section headed "Where do I find pricing for our legacy models?". It lists four rates for models the company no longer sells to anyone new. Command, the model that gave the family its name, is billed at one dollar per million input tokens and two per million output. Command-light costs 30 and 60 cents. Both run on a 4,000-token context window. The comparison with what Cohere sells today is not close. Command R 08-2024 handles 128,000 tokens — thirty-two times more — for 15 cents per million input tokens. That is one-sixth of what a customer still on Command is paying. Command R7B, the small model released in December 2024, costs under four cents per million input tokens, roughly one twenty-seventh of the Command rate, and also carries the 128,000-token window. What makes this durable is the absence of a deadline. Cohere's own lifecycle policy defines a deprecated model as one that is closed to new customers but stays available to existing users "until retirement", with a shutdown date "assigned at that time". For the five models deprecated in September 2025 — command, command-light, command-r-03-2024, command-r-plus-04-2024 and the Summarize endpoint — no such date has been published. The company has shown it will set one when it means to: the April 2026 retirement of Embed v2.0 and two Aya 8B models was announced with a firm date and named replacements. Deprecation also closed fine-tuning for the classic Command models, and previously fine-tuned versions stopped being accessible. So the customers who remain are running an unmodifiable 2022-generation model, on the shortest context window Cohere has ever sold, at the highest per-token price in the company's public price list — with nothing on the calendar to make them move. wujec.ai has added catalogue profiles for Command and Command-light, completing the classic Command line alongside the Command R and Command A families.
Cohere Command →xAI's newest Grok takes half the context of its March model and charges 140% more for output
There is a habit in this industry of reading a version number as a promise: the higher it is, the more the model does for the money. xAI's own price list breaks that habit. Grok 4.20 went live in March 2026 with a context window of one million tokens, priced at $1.25 per million input tokens and $2.50 per million output. Grok 4.3, which followed at the end of April, kept both numbers. Then the line turned. Grok 4.5 in July and Grok 4.6 in August each take 500,000 tokens - half of what their predecessors take - and are billed at $2.00 input and $6.00 output. Measured against the March model, the newest flagship costs 60% more to prompt and 140% more to answer, on a context window half the size. The two-tier billing that xAI applies across the line sharpens the gap. Any request whose prompt reaches 200,000 tokens is repriced in full at double rates, so a long document sent to Grok 4.6 is charged at $12 per million output tokens, against $5 for the same document sent to Grok 4.20. There is a second, quieter difference: Grok 4.20 and 4.3 accept Batch API requests at a 20% discount, and Grok 4.6 does not accept them at all. None of this is a deprecation story. xAI has not announced a shutdown date for either spring model, and both are still listed in the company's documentation and priced there. What the price list shows is that xAI's newer models are not successors in the usual sense - they are a separate, more expensive tier that trades context for whatever the company gained elsewhere, and the older, longer-context models remain the cheaper way to buy a million tokens from xAI. For comparison, the models xAI removed from that list did leave: Grok 3 and Grok 4 are no longer in the catalogue at all, and no shutdown date was ever published for them either.
Grok 4.20 →One model, three deaths: Claude 3 Haiku goes dark on Google Cloud tomorrow, four months after its maker switched it off
Anthropic retired claude-3-haiku-20240307 from its own API on 20 April 2026. The model has been running ever since — on somebody else's infrastructure. Google Cloud lists the same model as deprecated since 23 February 2026 and scheduled to be shut down on 23 August 2026: 125 days after the maker's own retirement date. Amazon Bedrock goes further still. There Claude 3 Haiku entered the Legacy state on 10 March 2026 and carries an end-of-life date of 10 September 2026 — 143 days past the maker. The Bedrock entry adds a wrinkle the other two do not have. Since 10 June 2026 the model has been in what AWS calls public extended access: a phase reserved for models that have already spent at least three months in Legacy, in which the provider may raise the price and AWS tells customers to expect exactly that. In other words, the last stretch of a model's life can be its most expensive. Anthropic states the rule plainly on its deprecation page: the dates published there apply to Anthropic-operated platforms only, and partner clouds set their own schedules. The gap runs in both directions — Google switched off Claude 3 Opus on 1 August 2025, months before the maker's own cut-off, while it kept Haiku alive four months longer. The practical consequence for anyone reading a model's retirement date: that date describes one platform, not the model. Claude 3 Haiku shipped in March 2024, has been unavailable from its maker since April, and will still be answering requests on Amazon's cloud three weeks from now.
Claude 3 Haiku →OpenAI took a third off its flagship's output price — and gave the cut an expiry date
On 21 August 2026 OpenAI's API changelog recorded a new price for GPT-5.6 Sol, the top tier of its mid-2026 frontier line: $4 per million input tokens and $20 per million output tokens, against $5 and $30 at launch. That is 20 percent off input and 33 percent off output. The pricing table moves with it — cached input drops from $0.50 to $0.40, with cache writes listed at $5. The wording is the part worth reading twice. OpenAI does not call this a new list price; it calls it promotional pricing, „available at least through November 21, 2026”. That is three months of guaranteed cover. Nothing in the changelog or on the pricing page says what happens on 22 November, and a permanent rate and a promotion that may lapse are different things to build a budget on. The cut also changes the shape of the line. Terra is unchanged at $2 and $12 per million tokens, Luna at $0.20 and $1.20. Until this week, reaching for the flagship instead of the middle tier cost two and a half times as much on both input and output. On the new rates it costs twice as much on input and roughly 1.7 times as much on output — by our arithmetic from the producer's own table, the smallest premium the top tier has carried since the line launched on 9 July 2026. What did not move is the cliff at the top of the context window. Any request whose input exceeds 272,000 tokens is still billed at twice the input rate and one and a half times the output rate for the whole request, which now works out at $8 and $30 — the long-context band therefore costs the same per output token as short-context calls did a week ago. The discounted queues follow the new rate: batch and flex both list Sol at $2 and $10, Fast mode at $8 and $40. The Daybreak Blue alias, which points approved defenders at gpt-5.6-sol, is priced identically.
GPT-5.6 Sol →Z.ai never prints how big its models are — its own weight files do, and the fifth generation is twice the fourth
Z.ai documents fifteen text models and six vision models on one price page, and describes none of them by size. The model cards list context windows, output limits, modalities and features; the parameter count, the detail every comparison starts from, is simply absent. Press coverage fills the gap with figures the company has never confirmed. It does not have to stay that way, because Z.ai publishes the weights. The safetensors index of each open model states the total exactly: GLM-4.6 (September 2025) and GLM-4.7 (December 2025) come to 356.8 and 358.3 billion parameters, both built as 92 layers with 160 routed experts plus one shared, eight active per token. The fifth generation doubles that. GLM-5, GLM-5.1 and GLM-5.2 all total roughly 753 billion parameters across 78 layers, with 256 routed experts plus one shared — the same skeleton reused three times, from February to June 2026. The price table follows the same split. Both fourth-generation models cost $0.60 per million input tokens and $2.20 per million output, and the line goes down to GLM-4.7-FlashX at $0.07/$0.40 and GLM-4.7-Flash, which is free. The fifth generation starts at $1.00/$3.20 for GLM-5 and settles at $1.40/$4.40 for GLM-5.1, GLM-5.2 and GLM-5.3 — 2.3 times the fourth-generation price for 2.1 times the parameters. One model breaks the pattern, and it is worth noting: GLM-5V-Turbo, the vision-and-coding model from April 2026, has no published weights at all. Z.ai repeatedly says it delivers its results "at a smaller model size" — and that is the one claim in the family which, for now, cannot be checked against a file.
GLM-5.1 →ByteDance's cloud switched on the DeepSeek price rise — and left the older build of the same model untouched at a quarter of the price
The increase BytePlus announced five days ago is now live. Since 00:00 Beijing time on 21 August 2026, ModelArk bills deepseek-v4-flash-ga-260731 at $0.44 per million input tokens, $1.32 per million output tokens and $0.014 for cache hits. Batch inference is charged at exactly half: $0.22 and $0.66. What the notice does not mention is the line directly below it in the same table. The April build of the same model, deepseek-v4-flash-260425, is still on sale and its rates were not touched: $0.14 in, $0.28 out. That is 3.1x less on input and 4.7x less on output than the version the platform now recommends. Only one number moves the other way — cache-hit input on the old build costs $0.028, twice the new rate. The rise puts ModelArk exactly level with DeepSeek's own list price, but only at one hour of the day. The vendor charges $0.44 / $1.32 during peak hours (01–04 and 06–10 UTC) and precisely half of that for the remaining eighteen hours. ModelArk has no time-of-day rate at all. The only way to halve the bill there is batch inference — the same 50% discount, paid for with latency instead of a clock. The parity is not limited to the small model. deepseek-v4-pro-ga-260813, the final August build of the larger model, is listed on ModelArk at $1.32 input and $3.96 output — again the vendor's peak rate to the cent, again with no off-peak equivalent.
DeepSeek-V4-Flash →OpenAI's model catalogue still advertises eight models that stopped answering — one of them 13 months ago
OpenAI's public catalogue of models for developers lists 96 entries. Eight of them are models that, according to the company's own deprecations table, no longer answer at all. The list is short enough to print. gpt-4.5-preview was switched off on 14 July 2025. o1-preview followed on 28 July 2025. o1-mini and both moderation entries — text-moderation-latest and text-moderation-stable, two catalogue rows for the same model — went dark on 27 October 2025. codex-mini-latest ended on 12 February 2026, chatgpt-4o-latest on 17 February 2026, and gpt-4-turbo-preview on 26 March 2026. The oldest of them has been sitting in the catalogue as a live product for thirteen months. Only one of the eight admits it. gpt-4.5-preview is described as a "deprecated large model"; the other seven read exactly like an offer. o1-preview is a "preview of our first o-series reasoning model", chatgpt-4o-latest is "the GPT-4o model used in ChatGPT" — present tense, for a model that has not been used in ChatGPT since February. Nothing in those descriptions distinguishes them from the models that work. The detail that turns an editorial slip into a practical problem is what the individual model pages still carry. The page for gpt-4-turbo-preview publishes a full commercial card: 10 dollars per million input tokens and 30 per million output, a rate-limit table for tiers 1 through 5, and a side-by-side comparison against GPT-4o mini and o3-mini. A developer who arrives there from a search engine — rather than from the deprecations page the catalogue mentions once in its preamble — has no signal that the model has been off since March. One of the eight also points at a replacement that does not exist under the name given. The deprecations table sends codex-mini-latest users to gpt-5-codex-mini; the catalogue has no such entry. What it does have is GPT-5.1-Codex mini, which is presumably the same intention with a different name — but the reader is left to make that leap. We checked two claims that looked like they belonged on this list and do not. babbage-002 and davinci-002 appear in the deprecations table with a 2024 date, but that row withdraws the ability to fine-tune them, not the models themselves; both still answer. The distinction matters here, because it is exactly the kind of ambiguity that makes a catalogue of 96 entries hard to audit — including for the company that publishes it.
Every pinnable ChatGPT model has now been switched off in OpenAI's API — what is left costs four times more
OpenAI has always sold two different things under similar names: the models in its API, and the model that actually answers inside ChatGPT. The second kind was reachable through a small family of aliases — chatgpt-4o-latest, then gpt-5-chat-latest and its successors. As of 10 August 2026, every one of them that carried a version number has been switched off. The sequence is short. chatgpt-4o-latest was removed on 17 February 2026. gpt-5-chat-latest and gpt-5.1-chat-latest were shut down together on 23 July 2026, under an announcement made on 22 April. gpt-5.2-chat-latest and gpt-5.3-chat-latest followed on 10 August 2026, deprecated on 8 May. In each case OpenAI named GPT-5.6 Sol — an API model, not a chat one — as the replacement. What remains is a single entry called chat-latest, and the wording on its card is the point: it "points to the latest Instant model currently used in ChatGPT" and "the underlying model snapshot will be regularly updated". There is no version number to pin. A developer who wants the behaviour of the ChatGPT model can still have it, but not frozen, and not with a guarantee that today's answers will resemble next month's. The price moved in the same direction. gpt-5-chat-latest and gpt-5.1-chat-latest cost USD 1.25 per million input tokens and USD 10 per million output. The 5.2 and 5.3 snapshots raised that to USD 1.75 and USD 14. chat-latest is listed at USD 5 and USD 30 — four times the input price of the first two, and precisely the price of the GPT-5.6 Sol flagship. Part of that buys real capability: the window grew from 128,000 tokens to 400,000, and the output ceiling from 16,384 tokens to 128,000, an eightfold increase. But the pinnable versions were the cheap ones, and they are the ones that are gone. There is also a rule behind the timing, and OpenAI publishes it. Its deprecation policy promises generally available models at least six months of notice, and specialised variants at least three — with "chat variants such as gpt-5.1-chat-latest" given as the first example of the latter. The line closest to the consumer product carries the shortest guarantee in the catalogue. One footnote for anyone reading the documentation directly: the card for GPT-5.1 Chat still describes it as the snapshot currently used in ChatGPT and invites developers to test chat improvements on it. The deprecation table has overtaken that sentence. And a distinction worth keeping straight — what ended was the API alias, not the model's presence in ChatGPT itself.
OpenAI Chat Latest →OpenAI told nine models to move here — and three of those destinations switch off on 28 September
OpenAI's deprecations page is the most orderly document the company publishes: a shutdown date, a model, and a recommended replacement, in 134 rows going back to 2023. Read down the third column and a pattern appears that no single row shows. The destinations are dying too. On 4 January 2024 the original GPT-3 base models were switched off. The table told `ada` and `babbage` customers to move to `babbage-002`, `curie` customers to `davinci-002`, and the six InstructGPT models — `text-ada-001` through `text-davinci-003` — to `gpt-3.5-turbo-instruct`. Those three destinations all shut down on the same day, 28 September 2026. Nine models were routed into them; everyone who complied packs twice in thirty-three months. It is not confined to the old base line. `gpt-4o` is the recommended replacement in twelve rows, covering the GPT-4 snapshots and the vision previews — and gpt-4o itself goes dark on 23 October 2026. `o1-preview` was pointed at `o3`, which ends on 11 December 2026. `o1-mini` was pointed at `o4-mini`, which ends on 23 October 2026, the same day as gpt-4o. The sharpest case runs backwards. `chatgpt-4o-latest` was deprecated on 18 November 2025 and removed on 17 February 2026, with `gpt-5.1-chat-latest` named as its replacement. That replacement was itself shut down on 23 July 2026 — five months after the model it was supposed to rescue. A developer who migrated on schedule had less than half a year before the same task landed again. None of this is hidden, and none of it is a broken promise: OpenAI's policy commits to notice periods, not to the longevity of whatever the table happens to name. The company also gave the 28 September cohort unusually long warning — the notice went out on 26 September 2025, more than twelve months ahead, against the six months a general model normally gets and the three months a preview gets. But the recommended replacement today is `gpt-5.6-terra`, at 2 dollars per million input tokens and 12 per million output. babbage-002 costs 40 cents on both sides. For a customer who only ever wanted raw text completion, the official route out is five times dearer on input and thirty times dearer on output — while GPT-5 nano, sitting in the same price list at 5 cents per million input tokens, is not mentioned in the table at all. The migration column names a default, not a fit. When babbage-002 and davinci-002 go, the last models descended directly from the GPT-3 base line leave the API with them — six years after that line established what a large language model was, and two migrations after their users were first told to move.
OpenAI davinci-002 →OpenAI still sells a 2022 embedding model at five times the price of its better replacement
Embedding models are the least glamorous part of an AI price list and the easiest to forget about once wired in. They turn text into vectors of numbers so that a system can measure how related two pieces of text are — the machinery behind search, clustering, recommendations and classification. Nobody demos them. They just sit in a pipeline and bill. Which is what makes the current OpenAI price list worth reading twice. **The same vector, one fifth of the money.** text-embedding-ada-002 shipped on 15 December 2022 and became the default of the early retrieval era. Its successors arrived on 25 January 2024. On OpenAI's own published figures, ada-002 scores 31.4% on MIRACL (multilingual retrieval) and 61.0% on MTEB (English); text-embedding-3-small scores 44.0% and 62.3%. Both produce 1,536-dimension vectors. The new one is better on both axes at the same output size — and is priced at $0.02 per million tokens against ada-002's $0.10. **The larger model makes the gap starker.** text-embedding-3-large reaches 54.9% on MIRACL and 64.6% on MTEB, at $0.13 per million tokens — three cents more than the model it replaced. And OpenAI notes something in the fine print of the 2024 announcement that is easy to miss: the new vectors can be truncated. Shortened to 256 dimensions, a text-embedding-3-large vector still beats a full 1,536-dimension ada-002 vector on MTEB. Six times smaller storage, better score, and still on the cheaper side of the comparison only if you count per token — per stored dimension it is not close. **Why the old model is still sold.** There is a real reason, and it is not sentiment. Embeddings are not portable between models: a vector from ada-002 and a vector from 3-small describe the same sentence in incompatible coordinate systems. Switching means re-embedding an entire corpus and rebuilding the index — for a large document store, a genuine engineering project. OpenAI has not announced a retirement date for ada-002, so nobody is being forced through that migration. The price difference is, in effect, what the vendor charges for not doing it. **A second, quieter case of a frozen price.** The same list holds TTS-1, OpenAI's first text-to-speech model. At launch on 6 November 2023 the company quoted $0.015 per 1,000 characters. Today it quotes $15 per million characters. Those are the same number: nearly three years, two generations of audio models, and no movement in either direction. **What this does not prove.** These are list prices for models that remain available, not evidence of anything hidden — OpenAI publishes all of it openly, and both older models still work exactly as documented. Nor is ada-002 a bad model; it was the best available when it shipped. The point is narrower and more practical: a default chosen in 2023 and never revisited is now the most expensive way to buy the weakest embeddings in the catalogue. wujec.ai has added profiles for all six models today, with the published benchmark figures side by side.
text-embedding-ada-002 →One OpenAI announcement, two deadlines: some models got six months to move, others three
On 22 April 2026 OpenAI published a single deprecation notice covering more than two dozen models. It carried two different shutdown dates. Fourteen models went dark on 23 July 2026. The rest run until 23 October 2026. Today, 20 August, sits neatly between the two — one half of that announcement is already history, the other half is still serving traffic. The split is not arbitrary, and OpenAI documents the rule. Generally available models get at least six months of notice. Specialised variants of those models get at least three. Preview models, identified by the word preview in their name, may be retired with as little as two weeks, and OpenAI explicitly advises against using them for business-critical work. What counts as a specialised variant is the interesting part. The company's own examples are chat variants such as gpt-5.1-chat-latest, Codex variants such as gpt-5.3-codex, and deep research variants such as o3-deep-research. On paper these are settings on a familiar base model. In practice they were sold as separate products, with their own model cards, their own pricing and their own API identifiers — and the July group included both deep research models, four Codex models and three chat models. The contrast within one announcement is sharp. o1-pro and o3-deep-research were deprecated on the same day. o1-pro, a generally available model, runs until 23 October. o3-deep-research, classed as a variant, was switched off on 23 July. Same notice, twice the runway for one of them. Two things are worth saying plainly. First, these are minimums, not promises of equal treatment: o1-mini was given six months in 2025 while o1-preview, announced the same day, got three. Second, none of this is hidden — the policy sits at the top of OpenAI's deprecations page, above the tables. The cost falls on whoever builds a product on a variant without reading it. wujec.ai has just added profiles for the six missing o-series reasoning models, including both deep research models and both pro-tier models, each with its announced shutdown date. They now appear in the catalogue's lifecycle timeline alongside the models that outlived them.
OpenAI o3-deep-research →