On 10 October Alibaba's cloud stops serving its rivals' models. DeepSeek, Kimi, GLM and MiniMax all get the same replacement: Qwen
Published: 9/10/2026 · Source: Alibaba Cloud — Model Studio service notices ↗
Model Studio, Alibaba Cloud's model platform, publishes its retirements as plain service notices, and six of them now converge on a single moment: 00:00 on 10 October 2026, Beijing time. After that hour the listed models stop answering; applications still calling them get nothing back.
The list of Alibaba's own casualties is unremarkable housekeeping. The April notice retires qwen-turbo, qwen-vl-max, qwen-vl-plus, qwq-plus and qvq-max, all of them 2025-era products. The June notice goes further up the range and takes qwen3-max with it, along with qwen3-max-preview, qwen3.6-max-preview, qwen3-vl-flash and qwen3-coder-plus; buyers are pointed at qwen3.7-max at $2.5 and $7.5 per million tokens, at qwen3.6-flash, or at qwen3.7-plus.
The interesting notice is the one about snapshots. Its table runs to roughly forty entries, and a large part of them are not Alibaba's models at all. DeepSeek-R1 and its 0528 revision, DeepSeek-V3, V3.1, V3.2 and the experimental V3.2-exp, the R1 distillations into Qwen 7B, 14B and 32B, Zhipu's glm-4.6 and glm-4.7, Moonshot-Kimi-K2-Instruct and kimi-k2-thinking, MiniMax-M2.1 — all of them are hosted third-party models that Alibaba sold access to, and all of them switch off on the same day. Against every one of those rows the recommended replacement column says the same thing: qwen3.7-plus, at $0.4 and $1.6 per million tokens below 256K of input, $1.2 and $4.8 above it.
It is worth being precise about what this is and is not. None of these models is being discontinued by its maker. DeepSeek, Zhipu, Moonshot and MiniMax all publish open weights, and their models remain available from their own APIs and from other clouds; what ends on 10 October is Alibaba's willingness to host them. For customers who picked a Chinese rival's model precisely because it was available inside Alibaba's console, however, the practical effect is the same as a shutdown: migrate, move cloud, or accept Qwen.
The direction of travel is easier to read from the replacement column than from the announcements themselves. A platform that spent 2025 advertising the breadth of its model catalogue is spending 2026 narrowing it to the models it makes. Speech is going the same way — a separate notice, also dated 10 October, closes Alibaba's hosted qwen3-asr-flash and qwen3-tts-flash endpoints and sends users to fun-asr and cosyvoice, which are not Qwen models either.
Anyone running production traffic through Model Studio has under a month to check which of these names appears in their logs. The catalogue keeps profiles of the affected third-party models unchanged, because the models themselves are alive; the entry that ends is the one on Alibaba's price list.