News

What's happening in robotics and AI — curated by the wujec.ai editors.

Business8/22/2026 · MiniMax — pliki licencji przy wagach modeli na Hugging Face, dziennik wydań i cennik platformy MiniMax

MiniMax's most downloaded model is the one you may not sell anything with

MiniMax M2.7 is pulled from Hugging Face roughly 909,000 times a month — more than any other text model the Chinese lab has published, and more than four times the traffic of its own successor M3. It is also the only one in the family that forbids commercial use. The licence file shipped with the weights is titled "Non-Commercial License". Personal use, self-hosted deployment, research and experimentation are expressly free of charge. Everything else is not: selling a product or service that relies on the model, putting it behind a paid API, or deploying a fine-tuned derivative for commercial purposes all require prior written authorisation from MiniMax, requested by e-mail. Anyone who obtains that permission must also display "Built with MiniMax M2.7" on the product. That term is an outlier rather than a direction of travel. MiniMax M2 (October 2025) and M2.1 (December 2025) were released under MIT with a single added clause: a commercial product built on the model must show the model name in its interface. In M2 that duty starts only above 100 million monthly active users or $30 million in annual recurring revenue; in M2.1 it applies with no threshold at all. M2.5 (February 2026) moved to the company's own MiniMax Model License, which still allows redistribution but requires a fixed attribution notice to accompany every copy. M2.7, released a month later, closed commercial use altogether. M3, in June 2026, reopened it under a community licence that permits commercial deployment in exchange for visible attribution. The practical consequence is easy to miss, because nothing about the download page signals it. Four models in this line share the same body — 62 layers, 256 experts, eight routed per token, and 228,689,764,864 parameters in M2, M2.1 and M2.7, to the byte. A team that swaps one for another as a drop-in upgrade changes its legal position without changing a line of code. The strongest freely reusable model in the line is M2.5; the strongest one overall, by the producer's own benchmark table, is the one that needs a signature. wujec.ai has today added catalogue profiles for M2.1, M2.5 and M2.7, each stating the licence terms in full.

MiniMax M2.7
Releases8/17/2026 · Hugging Face (meta-models)

Meta promised its flagship weights a week ago — its download page still holds only the distilled model

On 10 August 2026 Mark Zuckerberg published a long essay arguing that American labs should lead the open-weight movement, and said Meta would open the weights of Muse Spark 1.2, its most capable model. The word he used for the timing was "soon". No date was given. Seven days later the company's model account on Hugging Face holds four repositories, and all four are the same smaller model. Muse-Glimmer-30B, the base weights, went up on 9 August. Next to it sit a GGUF conversion for local runtimes, a draft head for speculative decoding published as an "assistant" repository, and an ExecuTorch build for mobile deployment. There is no Muse Spark repository of any kind — not the 1.2 release, not the 1.1 one that preceded it in July. The distinction matters more than it may look, because the two models are not alternatives. Meta describes Muse Glimmer as a distillation of Muse Spark: a roughly 29.6-billion-parameter student trained on the teacher's outputs, with a 131,072-token context window against the flagship's million. What can be downloaded today is the compressed derivative of the model whose weights were promised, not the model itself. Demand for the derivative has been considerable. Meta's own four repositories record 780,323 downloads in the eight days since publication — 334,099 for the base weights, 395,175 for the GGUF conversion, 43,909 for the drafter and 7,140 for the mobile build. A single community quantisation published by Unsloth adds another 755,125, taking the family past 1.5 million downloads in little over a week. All of it carries a plain Apache 2.0 licence with no revenue or user thresholds, which is itself a departure from the bespoke community licences Meta attached to its earlier open releases. Until the flagship weights appear, the practical position is unchanged: Muse Spark 1.2 remains reachable only through Meta's interfaces, and the open-weight commitment stands as an announcement rather than a file. wujec.ai will update the Muse Spark 1.2 profile on the day a repository appears.

Muse Spark 1.2
Research8/15/2026 · Z.ai — blog premierowy GLM-5.3 i rejestr Z.ai Security Disclosure Ledger

A model looked at 269 open-source projects and found 2,436 flaws — the oldest dating to 1981

Z.ai released GLM-5.3 on 14 August and buried the most interesting number deep in the announcement. Working with security teams in China, the company pointed the model at real open-source codebases. After expert review, screening and deduplication, it had identified **2,436 vulnerabilities across 269 projects** — 107 rated critical, 990 high, 1,286 medium and 53 low. The findings span kernels, operating systems, browser engines, infrastructure libraries, web applications and network protocols. The striking part is not the count but the age. By Z.ai's figures the average flaw had sat in its codebase for **26.6 years** before anyone noticed, and the oldest was introduced in **1981** — forty-five years of impact. These are not fresh regressions in fast-moving projects; they are defects that survived every human code review, static analyser and fuzzing campaign of the last four decades. Z.ai says the capability was not the goal. Vulnerability-discovery environments were added to the post-training mix expecting the model to get better at spotting isolated flaws; what emerged, in the company's words, was a model that reasons across multiple stages of exploitation and forms coherent plans for complete chains. The benchmark numbers back a narrower claim: on CyberGym, which starts from source code and tests whether a model can find and validate a vulnerability, GLM-5.3 scores 84.5%, up from 77.2% and marginally ahead of Claude Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). Further along the chain the lead vanishes — on ExploitBench it reaches 54.4% against 78.0% for Mythos 5, and on ExploitGym it completes 105 tasks in two hours against 181. Z.ai states this plainly: the advantage sits at the front of the exploitation chain, and the gap to the closed frontier is widest where the capability matters most. What makes the disclosure unusual is that it comes with a paper trail. Z.ai has published a public ledger at cvd.z.ai recording each finding as it moves through coordinated disclosure: affected project, severity, CVE where assigned, and how long the flaw had been in the code. As of the announcement, **53 findings are public and 2,383 remain under embargo** — which is itself the story. A single model run has produced a backlog of undisclosed vulnerabilities larger than most national CERTs handle in a year, and the maintainers of those 269 projects now hold the timetable. GLM-5.3 was not downloadable at launch. Z.ai promised weights about two weeks after the announcement and lists the API as coming soon; for now the model runs through the GLM Coding Plan subscription and the ZCode agent. When the weights do land, the same capability that filled that ledger becomes available to anyone with the hardware to run it — which is the argument for staged release, and the argument against it, depending on who is making it. *Editorial note: wujec.ai reports on security capability as a published property of these models. We do not reproduce exploit material and we link only to the vendor's own coordinated-disclosure record.*

GLM-5.3
Business8/15/2026 · Anthropic — dokumentacja cenowa

The 50% rise is off: Anthropic makes Claude Sonnet 5's $2/$10 the standard price

Sixteen days before the deadline, Anthropic has cancelled the price increase it scheduled for Claude Sonnet 5. The company's pricing documentation now states plainly that the $2 per million input tokens and $10 per million output tokens, announced at launch as an introductory rate running through 31 August 2026, is the standard price, and that the increase to $3 and $15 planned for 1 September will not occur. This reverses the situation we reported on 6 August, when the same documentation carried the expiry date as a footnote. What was a discount with a countdown is now simply the price. Sonnet 5 keeps its position in Anthropic's ladder — Fable 5 at $10 / $50, Opus 5 at $5 / $25, Sonnet 5 at $2 / $10, Haiku 4.5 at $1 / $5 — but it is no longer the only current Claude model sold below its own list. One qualification from that earlier piece still holds, because it was never about the rate card. Sonnet 5 uses the tokenizer introduced with Opus 4.7, and Anthropic's own note says the same text yields roughly 30% more tokens on it than on the previous generation; Sonnet 4.6 and earlier use the older tokenizer. On the headline, Sonnet 5 is a third cheaper than Sonnet 4.6's $3 / $15. Measured per page of text rather than per token, the saving is closer to a tenth. The same arithmetic applies to the context window: the company puts one million tokens on Sonnet 5 at about 555,000 words, against roughly 750,000 words for the same million on Sonnet 4.6. The timing invites one comparison. Two days ago Google presented Gemini 3.7 Flash at $0.75 and $3.75 per million tokens, described as half the price of its predecessor — a rate its own price list dates to 31 December 2026, after which it doubles. Within one week, then, two vendors have taken opposite decisions about the same instrument: one has removed the expiry date from a discount, the other has left it in place. For anyone budgeting a year ahead, that difference matters more than the headline figures. Our Claude Sonnet 5 profile has been updated with the standard price and the cancelled increase.

Claude Sonnet 5
Releases8/13/2026 · DeepSeek (Hugging Face model card)

DeepSeek ships the final V4-Pro: open weights, 1M context, agentic scores up sharply

DeepSeek published DeepSeek-V4-Pro-0813 on Hugging Face on 13 August 2026, describing it as the official release of V4-Pro and the successor to the preview weights that had carried the name since April. The repository went up under the MIT licence, so the full model — 1.65 trillion parameters in a sparse mixture-of-experts layout, 61 layers, 384 routed experts with six active per token, shipped in FP8 — can be downloaded and self-hosted. The context window stays at 1,048,576 tokens. The architecture is unchanged from the preview; what is new is a DSpark speculative-decoding module attached to the model and a reworked reasoning control. The `reasoning_effort` parameter now takes three levels — low, high and max — letting callers trade deliberation time against cost. DeepSeek also dropped the Jinja chat template in favour of a documented Python encoder, a change that will require work from anyone running the weights locally. The gains DeepSeek reports are concentrated in agentic and coding work, and they are large. On the company's own table the new build reaches 62.7% on DeepSWE against 12.8% for the preview, 61.5% on NL2Repo against 38.5%, 83.3% on Cybergym against 52.7% and 87.9 on Terminal Bench 2.1 against 72.1. On Humanity's Last Exam it scores 42.7% without tools and 60.0% with them, up from 37.7%. Against rivals in the same table the picture is closer. Kimi K3 edges it on Terminal Bench (88.3) and DeepSWE (67.5), and Anthropic's Fable 5 leads on Humanity's Last Exam (53.3 without tools). What DeepSeek keeps is the position it has held all year: comparable results at frontier level, with the weights published rather than rented. Benchmark figures here are the vendor's own and have not yet been reproduced independently.

DeepSeek-V4-Pro
Releases8/12/2026 · OpenAI

OpenAI ships a model trained to stop refusing: GPT-5.6 Cyber answers 95% of hacking prompts, and almost nobody can buy it

OpenAI announced GPT-5.6 Cyber on 10 August 2026, and the headline figure is unusual: it is not a capability score but a compliance rate. On the company's own Advanced Cybersecurity Completion Rate — how often a model answers prompts about exploit chains, authentication bypass and privilege escalation rather than declining — the new model responds to 95.0% of them. GPT-5.6 Sol behind standard guardrails answers 1.5%. Last year's GPT-5.5 Cyber managed 57.3%. The model is built on GPT-5.6 Sol and further trained for zero-day discovery and exploit development. Its usefulness has already been demonstrated on live software: OpenAI used it to study V8, Chrome's JavaScript engine, and found two previously unknown bugs that chain into an escape from the heap sandbox. Google patched them as CVE-2026-15903. OpenAI also reports at least five vulnerabilities in a widely used mobile operating system, three critical flaws in a popular database, and more than 400 privilege-escalation issues in an OS kernel — all now in coordinated disclosure. What makes the release notable is not that the model is stronger, because in places it is not. On ExploitBench 3, a harder V8 task with sandbox protections left on, ordinary GPT-5.6 Sol solves more within the standard 300-turn budget; the gap only narrows at 600 turns. On OpenAI's internal vulnerability-report evaluation, Cyber scores below Sol, which the company blames on shorter, thinner write-ups. Under the Preparedness Framework it lands at High for cyber capability — the same rating as Sol — and below the Critical threshold. What changed is who gets to ask. GPT-5.6 Cyber exists only inside Daybreak Red, the offensive tier of an access programme OpenAI expanded the same day; Daybreak Blue, the defensive tier, keeps the general-purpose models with guardrails tuned for defence. Entry requires identity verification, approved-use restrictions, monitoring and legal attestations, and from 1 September 2026 every individual Daybreak account must use a hardware security key. Early partners named by OpenAI include Accenture, IBM, Capgemini, EY, KPMG, PwC, Palo Alto Networks, CrowdStrike, Cloudflare, Akamai, Fortinet, Sophos and SpecterOps. The technical shape is narrower than Sol's: a 400,000-token context window against Sol's 1,050,000, output up to 128,000 tokens, text and image in, text out, knowledge to 16 February 2026. It runs on the Responses endpoint only, with no chat completions, batch or fine-tuning, and lists at USD 12.50 per million input tokens and USD 75 per million output — two and a half times Sol's price. A full system card has been promised at a later date.

GPT-5.6 Cyber
Business8/6/2026 · Anthropic — dokumentacja modeli

Claude Sonnet 5's introductory price expires on 31 August — and the bill rises 50%

**Update, 15 August 2026:** Anthropic has cancelled this increase. Its pricing documentation now names $2 / $10 the standard price for Claude Sonnet 5 and states that the rise to $3 / $15 on 1 September will not occur. See: The 50% rise is off (/news/claude-sonnet-5-price-rise-cancelled-2-10-permanent). The paragraph below on the tokenizer still applies. Anthropic's model documentation carries a footnote that is easy to miss and expensive to ignore: the introductory pricing of $2 per million input tokens and $10 per million output tokens applies to Claude Sonnet 5 only through 31 August 2026. From 1 September the model reverts to its list price of $3 and $15 — a 50% increase on both sides of the meter, arriving in under four weeks. Sonnet 5 is not a marginal product for Anthropic. Released on 30 June 2026, it is the default model for Claude's free and Pro users and the company's declared best combination of speed and intelligence, with a native one-million-token context window, 128,000 output tokens on the Messages API and 63.2% on SWE-Bench Pro. It is also, as of today, the only current Claude model sold below its own list price: Fable 5 stands at $10 / $50, Opus 5 at $5 / $25 and Haiku 4.5 at $1 / $5, all at list. There is a second, quieter movement in the same direction, and it has nothing to do with the price card. Anthropic's own documentation puts the one-million-token window of Fable 5, Opus 5 and Sonnet 5 at roughly 555,000 words, while the same one million tokens on Opus 4.6 and Sonnet 4.6 held about 750,000 words. For Fable 5 the company states the mechanism outright: it uses the tokenizer introduced with Opus 4.7, and the same text produces roughly 30% more tokens than on the older models. Since tokens are the billing unit, an unchanged workload on the newer generation is metered higher before any rate change is applied at all. What this does not appear to be is a repricing of the frontier. Anthropic has not announced a change to any other model's rates, and the retirement floor for Sonnet 5 remains unmoved at 30 June 2027. The plain reading is that a launch discount is simply running out on schedule. For anyone budgeting against Sonnet 5, the two dates that matter are 31 August, when the discount ends, and the moment a workload is ported from a 4.6-generation model, when the token count itself changes. Our Claude Sonnet 5 profile now carries both the introductory and the list price.

Claude Sonnet 5
Releases8/6/2026 · Perplexity — dokumentacja deweloperska

Perplexity's own model steps aside: the Sonar API becomes a shop selling OpenAI, Anthropic and Google

Perplexity's developer documentation now opens with a line that reads like a quiet change of business: "Sonar Chat Completions is now Agent API." The interface the company launched on 21 January 2025 to sell its own search-grounded model is being succeeded by one built around an agent loop — and around other companies' models. The Agent API is described as a unified interface for agent applications, with built-in tools the old endpoint never had: web search, URL fetching, code sandboxes, MCP servers, people search and finance search. Instead of a `messages` array with `choices`, a request produces a typed `output` array recording each step the agent took — reason, act, observe, continue, all inside a single call. The more interesting detail is the model list. Alongside Sonar, the Agent API brokers model families from OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI and NVIDIA. A platform that existed to sell one proprietary answering model is now largely a way to buy everyone else's, with Perplexity's search stack as the wrapper. Sonar itself has not been switched off. The documentation says Sonar Chat Completions remains supported, while calling the Agent API more performant and cost-effective for production workloads — the standard phrasing of a migration that has not yet been given a deadline. Sonar Pro pricing is unchanged at 3 USD per million input tokens and 15 USD per million output, plus 6-14 USD per 1,000 requests depending on how much searching the answer required. The strategic reading is straightforward. Sonar's advantage was never raw reasoning; it was reading the live web and citing it, which is why it leads factuality benchmarks like SimpleQA while trailing frontier models elsewhere. Wrapping that retrieval layer around whichever model a developer prefers is a better use of the asset than asking them to accept Perplexity's model as well.

Perplexity Sonar Pro
Releases8/3/2026 · Forbes

Alibaba launches Qwen3.8-Max, a 2.4-trillion-parameter flagship

Alibaba released Qwen3.8-Max on 3 August 2026, the largest model the Qwen family has produced so far. It is a sparse mixture-of-experts design: 2.4 trillion parameters in total, of which roughly 95 billion are activated per token, with a context window of up to one million tokens. The model is available worldwide through Alibaba Cloud's Model Studio APIs and through QwenWork, the company's workplace agent platform. Alibaba said full model weights would follow for public download, alongside a smaller Qwen3.8-27B variant aimed at hardware-constrained deployments. On the Arena.AI leaderboard Qwen3.8-Max became the highest-ranked Chinese model for text tasks and placed second globally for vision, putting it in the same bracket as current frontier systems from OpenAI and Anthropic. Alibaba shares rose sharply in Hong Kong on the announcement.

Qwen3.8-Max
Releases7/21/2026 · Google DeepMind

Google releases Gemini 3.6 Flash across Search, Android and Workspace

Gemini 3.6 Flash, released July 21, 2026, continues Google's strategy of fast, production-friendly frontier models woven into its products — Search AI Mode, Android assistants and Workspace. It balances low latency with strong multimodal reasoning, and powers robotics work through the Gemini Robotics line.

Gemini 3.6 Flash
Releases7/19/2026 · AIToolsReview

DeepSeek's 1.6-trillion-parameter flagship goes GA — and starts charging by the clock

DeepSeek-V4-Pro left preview on 19 July 2026, three months after its first public release. The weights are the same ones published on 24 April: a sparse mixture-of-experts model with 1.6 trillion total parameters, roughly 49 billion of them active per token, a one-million-token context window and an MIT licence that allows anyone to download and self-host it. What general availability changed was the commercial side. DeepSeek introduced peak-time pricing for the first time, doubling rates during Beijing business hours; off-peak the API costs 0.435 US dollars per million input tokens and 0.87 per million output. Cached input is charged at roughly a hundredth of the off-peak input rate. On 24 July the company retired the legacy deepseek-chat and deepseek-reasoner endpoints, so integrations now have to name a model explicitly. On DeepSeek's own GA figures the model scores 80.6% on SWE-bench Verified — the strongest published result for an open-weights model — with a Codeforces rating of 3,206, 57.9% on SimpleQA Verified and 37.7% on Humanity's Last Exam. Independent measurement tells a subtler story: Artificial Analysis rated V4-Pro at 44 on its Intelligence Index at GA, down from 52 at preview. The model did not get worse; the field moved during the three months it spent in preview. The smaller DeepSeek-V4-Flash and the earlier DeepSeek-R1 remain available, and all three now have profiles in the wujec.ai catalogue.

DeepSeek-V4-Pro
Releases6/30/2026 · Anthropic

Anthropic ships Claude Sonnet 5 with a native 1M-token context

Released on June 30, 2026, Claude Sonnet 5 is Anthropic's workhorse frontier model: 63.2% on SWE-Bench Pro, a native one-million-token context window and strong tool-use reliability at Sonnet pricing. It became the default model for Claude's free and Pro users, bringing frontier-grade agentic coding to the mainstream.

Claude Sonnet 5