All newsBusiness

Stripe paid over $7bn for a shop where one model carries 31 prices — the company that built it ranks 19th

Published: 8/18/2026 · Source: OpenRouter API (models/endpoints), Z.ai pricing documentation, Bloomberg

Bloomberg reported on 16 August 2026 that Stripe had clinched a deal to buy OpenRouter for more than $7 billion — 5.4 times the $1.3 billion the company was valued at in its May funding round. OpenRouter does not train models. It resells access to other people's, and publishes its full catalogue through a public API. We queried that API on 18 August 2026 and counted what is actually on the shelves. **One model, thirty-one prices.** GLM-5.2, the flagship open-weight model from Chinese lab Z.ai, is served by 31 separate endpoints. The cheapest is Novita at $0.489 per million input tokens; the dearest is Alibaba at $2.31 — a spread of 4.7x for the same weights. Z.ai's own endpoint charges $1.40 / $4.40, exactly the figure in its published price list, and lands nineteenth out of thirty-one. The cheapest reseller undercuts the model's author by 65%. **The pattern is not one lab's oddity.** Across the open-weight models we checked, the maker's position among its own hosts is close to random: | model | hosts | price spread | where the maker ranks | |---|---|---|---| | GLM-5.2 (Z.ai) | 31 | 4.73x | 19th — 2.87x the cheapest | | gpt-oss-120b (OpenAI) | 20 | 11.67x | absent — OpenAI does not host it | | DeepSeek V4 Pro | 18 | 2.89x | 1st — nobody undercuts it | | Kimi K2.7 Code (Moonshot AI) | 15 | 2.84x | 10th — 1.42x the cheapest | | MiniMax M3 | 12 | 3.26x | 6th — 1.30x the cheapest | | Qwen3.8 2.4T A95B (Alibaba) | 7 | 1.25x | joint cheapest on input | The widest gap belongs to a model whose maker sells no hosting at all: gpt-oss-120b runs from $0.03 per million at CoreWeave to $0.35 at Cerebras — an eleven-fold range with no official price to anchor it. **A batch mode that costs more than the live one.** OpenRouter lists 61 batch variants of models on its books. Sixty of them behave as expected: batch costs exactly half the standard rate — the same 50% discount at OpenAI, Anthropic, Google and NVIDIA alike, to four decimal places. One does not. GLM-5.2's batch tier costs $0.70 / $2.20 against $0.49 / $1.54 for live requests — 43% *more* — and comes with a 512,000-token window instead of 1,048,576. The reason is visible in the endpoint data: the batch tier has a single host, Together, pricing at half of Z.ai's official rate, while the live tier is routed to sellers already charging a third of it. Half of the list price still loses to the street price. **What these numbers do not prove.** They are not evidence that eighteen companies serve GLM-5.2 better than Z.ai does. Cheap endpoints frequently run reduced precision: the three cheapest MiniMax M3 offers are fp4, the cheapest Kimi K2.7 Code endpoint is int4, and quantisation changes what the model outputs. Context windows differ just as sharply — GLM-5.2 endpoints range from 96,890 tokens to 1,048,576, a tenfold gap, and the 96,890-token seller is not the cheapest one. Throughput, rate limits and data retention are not in these figures at all. And prices on a marketplace move: this is a single reading, taken on 18 August 2026. What the numbers do show is what Stripe is buying. Not a model, and not a discount — a price list that its own suppliers cannot see the bottom of.