News

What's happening in robotics and AI — curated by the wujec.ai editors.

Releases8/26/2026 · OpenAI — developer documentation (model cards and pricing)

OpenAI still calls a February model its best coder — and it knows nothing written after August 2025

OpenAI's documentation describes GPT-5.3-Codex as "the most capable agentic coding model to date". It went on sale on 24 February 2026. Six months later that sentence still stands, because nothing has replaced it: the dedicated Codex line has not had a new member since. The general line has not stood still in the same period. GPT-5.5 arrived on 24 April and the three-model GPT-5.6 family — Sol, Terra and Luna — on 9 July, and OpenAI's own catalogue entry for GPT-5.6 Sol now reads "start here for complex reasoning and coding". The specialist and the generalist are being pointed at the same job. The gap between them is not marketing. GPT-5.6 Sol carries a 1,050,000-token context window against the Codex model's 400,000, and a knowledge cutoff of 16 February 2026 against 31 August 2025. For a coding model the second number matters more than it looks: a cutoff of August 2025 means the tool writing your dependency call has never seen a year of releases, deprecations and breaking changes in the libraries it is calling. What the older model keeps is price. GPT-5.3-Codex costs $1.75 per million input tokens and $14 per million output, against $4 and $20 for GPT-5.6 Sol — a bit over half the input rate. It is also narrower by design: the Codex line runs only on the Responses API, with no Chat Completions, no batch, no fine-tuning. That leaves a choice OpenAI does not spell out anywhere in one place. The cheap specialist is frozen in time; the expensive generalist is current. Nothing in the documentation says the Codex line has been retired, and no shutdown date has been published for it — which is exactly why the six-month silence is worth noticing rather than assuming. Figures in this article come from OpenAI's own model cards and pricing tables, read on 26 August 2026; the release dates were cross-checked against an independent model registry, because OpenAI does not date its model cards.

Releases8/17/2026 · Qwen — karta modelu Qwen3.8-2.4T-A95B na Hugging Face

Alibaba's downloadable flagship is not the flagship it rents — the model card lists four things it cannot do

Qwen3.8-Max, the 2.446-trillion-parameter model Alibaba published for download on 8 August, is the first Max-class model the company has ever released as open weights. What almost no coverage mentions is that the file you download and the model you call through the API are not the same product — and the difference is written by Alibaba itself, in a note near the top of the model card. The note says the API version is "based on" the released checkpoint "with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools". Read as a list of what the download lacks, it comes to four items: - **No vision.** The repository ships a text-only architecture (`qwen3_5_moe_text`, pipeline tag `text-generation`). The hosted Qwen3.8-Max accepts images and video; the download does not. - **No way to switch reasoning off.** The open checkpoint always reasons before answering — depth is adjustable through `reasoning_effort`, but the non-thinking mode exists only in the API. - **262,144 tokens, not a million.** The card states the context length plainly: 262,144 natively, extensible to 1,010,000. The hosted version has a million by default; self-hosters have to extend it themselves and accept whatever that does to quality. - **No built-in tools.** Web search, code execution and the rest of the hosted toolchain are part of the service, not the weights. None of this is concealed — it is one paragraph in Alibaba's own card, which is precisely why it is worth repeating: the benchmark numbers circulating for Qwen3.8-Max (Terminal-Bench 2.1 at 86.6, SWE-bench Pro at 67.7, DeepSWE 1.1 at 56.6) were measured on the hosted configuration, and a self-hosted deployment starts from a narrower one. The card also settles what the model actually is architecturally, which the launch materials left vague. It is a hybrid: 92 layers arranged as 23 repetitions of three Gated DeltaNet blocks followed by one gated attention block, so only a quarter of the layers use full attention and the rest use linear attention. The mixture of experts holds 512 experts, of which 10 routed plus 1 shared fire per token. That is the design that makes 2.4 trillion parameters cost 95 billion per token to run — and it is the first time Qwen has shipped it at flagship scale. Our profile of the model now carries the native context figure, the hybrid architecture and an explicit note that the published weights are text-only.

Qwen3.8-Max →