News

What's happening in robotics and AI — curated by the wujec.ai editors.

Research9/13/2026 · Z.ai — GLM-5.3 model guide

Z.ai measures its model against a rival in two hours — but the two hours are set separately for each model

The most honest sentence in Z.ai's GLM-5.3 announcement is also the one that makes its own headline number hard to read: on the ExploitGym benchmark, the time budget every model gets is normalised model by model, using throughput figures the company leaves to the footnotes. The comparison reads as "105 tasks in two hours against 181" — but the two hours are not the same two hours for each contestant. ExploitGym counts how many exploitation tasks a model completes under a fixed time budget. Z.ai reports GLM-5.3 at 105 tasks in two hours and 130 in six, up from 29 and 39 for GLM-5.2. Only one closed model appears in that row: Mythos 5, at 181 and 247. A benchmark scored in wall-clock time measures two things at once — how well a model reasons and how fast it is served — and normalising the budget is a defensible way to separate them. It also means the headline gap depends on a throughput assumption the reader never sees. Two more qualifications sit in the same announcement, both of them stated by Z.ai and both routinely dropped when the numbers travel. The state-of-the-art claim on Terminal-Bench 3.0 and Agents' Last Exam is expressly a record among open-weight models, not against the field. And on the company's in-house Z.ai Code Bench, GLM-5.3 reaches 34.5% at maximum effort — ahead of Claude Opus 4.8 at the high setting, behind Claude Fable 5 at 39.5%. The summary at the top of the page mentions the first comparison and not the second. None of this makes the security result small. GLM-5.3 climbs from 77.2% to 84.5% on CyberGym, the best figure in the company's table, and Z.ai says the model found 2,436 vulnerabilities across 269 real projects, 1,097 of them medium-to-high severity. The company's own reading is the useful one: capability grows fastest exactly where the distance to the closed frontier is greatest. We have corrected our GLM-5.3 profile accordingly — the earlier version attributed the ExploitGym figures 181 and 247 to two different closed models, when both belong to one, at two different budgets.

GLM-5.3 →
Research9/4/2026 · Hugging Face (zai-org account)

Two thirds of Z.ai's model repositories are older than 2025. One of them carries most of what is left of their traffic.

Z.ai keeps 154 model repositories on Hugging Face. Ninety-eight of them - just under two thirds - were published before 2025, and together they were downloaded 810,718 times in the last thirty days. That is eight per cent of the account's 10.08 million downloads. The fourteen repositories created in 2026 took 8.38 million, or eighty-three per cent. The surprise is how concentrated the remainder is. ChatGLM2-6B, published in June 2023, accounts for 435,810 of those downloads on its own - more than half of everything the pre-2025 archive collects, and more than the vendor's current GLM-5, GLM-5.1, GLM-5.3 or GLM-4.7 individually pull in the same period. A three-year-old six-billion-parameter chat model outranks four of the company's own flagships on the download counter. Size is the readable explanation. ChatGLM2-6B runs on a 6 GB graphics card at INT4 quantisation and holds about 8,000 tokens of conversation there; the 2026 flagships do not fit consumer hardware at all. The rest is inertia: tutorials, university course material and fine-tuning recipes written in 2023 still point at that repository by name. One caveat belongs with the numbers. Hugging Face reports downloads over a rolling thirty-day window only, so these figures describe present-day pulls, not lifetime totals - a model released last week and a model released three years ago are measured on the same month. wujec.ai has added profiles for ChatGLM-6B, ChatGLM2-6B and GLM-4-9B-Chat, the three models this history runs through.

ChatGLM2-6B →
Releases8/28/2026 · Z.AI Developer Documentation / Hugging Face

Z.ai opened the weights of its cheap model — the flagship it promised two weeks ago is still closed

Z.ai published GLM-5.3-Flash on 26 August 2026 and put the weights on Hugging Face the same week, under a plain MIT licence. Two weeks earlier the company launched its flagship GLM-5.3 with a promise that its weights would follow in about a fortnight. That deadline has now passed, and the flagship repository does not exist: the newest Z.ai model anyone can download is the cheap one. The last flagship whose weights were actually released is GLM-5.2 from June 2026, at 753.3 billion parameters. GLM-5.3-Flash is less than half that size — 320 billion parameters with 18 billion active per token, and the published safetensors index confirms it at 321.3 billion. The smaller model is not merely a trimmed version. Z.ai describes it as the first open-weight frontier model to combine sparse attention with linear attention, and quantifies the gain against GLM-5.3: attention computation down 3.01 times, KV cache down 4.44 times. It is also the first natively multimodal model in the family, meaning it can look at a rendered interface and correct its own code rather than working blind. The price is where that architecture shows. GLM-5.3-Flash lists at 0.15 dollars per million input tokens and 0.50 per million output; GLM-5.3 costs 1.40 and 4.40. That is 9.3 times cheaper on input and 8.8 times cheaper on output — close to, but not quite, the one tenth of the price the company claims. A 50 percent promotion runs until 9 September 2026, which for now doubles the gap again. The performance claims stay in the maker’s own hands. Z.ai says the model beats GLM-5.2 across benchmarks and approaches Claude Opus 4.8 on coding and agentic tests, but publishes those results only as a chart image, and no independent laboratory has repeated them. What can be verified is the licence: the LICENSE file in the repository is the bare MIT text, with no attribution rider and no restriction on commercial use.

GLM-5.3-Flash →
Research8/15/2026 · Z.ai — blog premierowy GLM-5.3 i rejestr Z.ai Security Disclosure Ledger

A model looked at 269 open-source projects and found 2,436 flaws — the oldest dating to 1981

Z.ai released GLM-5.3 on 14 August and buried the most interesting number deep in the announcement. Working with security teams in China, the company pointed the model at real open-source codebases. After expert review, screening and deduplication, it had identified **2,436 vulnerabilities across 269 projects** — 107 rated critical, 990 high, 1,286 medium and 53 low. The findings span kernels, operating systems, browser engines, infrastructure libraries, web applications and network protocols. The striking part is not the count but the age. By Z.ai's figures the average flaw had sat in its codebase for **26.6 years** before anyone noticed, and the oldest was introduced in **1981** — forty-five years of impact. These are not fresh regressions in fast-moving projects; they are defects that survived every human code review, static analyser and fuzzing campaign of the last four decades. Z.ai says the capability was not the goal. Vulnerability-discovery environments were added to the post-training mix expecting the model to get better at spotting isolated flaws; what emerged, in the company's words, was a model that reasons across multiple stages of exploitation and forms coherent plans for complete chains. The benchmark numbers back a narrower claim: on CyberGym, which starts from source code and tests whether a model can find and validate a vulnerability, GLM-5.3 scores 84.5%, up from 77.2% and marginally ahead of Claude Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). Further along the chain the lead vanishes — on ExploitBench it reaches 54.4% against 78.0% for Mythos 5, and on ExploitGym it completes 105 tasks in two hours against 181. Z.ai states this plainly: the advantage sits at the front of the exploitation chain, and the gap to the closed frontier is widest where the capability matters most. What makes the disclosure unusual is that it comes with a paper trail. Z.ai has published a public ledger at cvd.z.ai recording each finding as it moves through coordinated disclosure: affected project, severity, CVE where assigned, and how long the flaw had been in the code. As of the announcement, **53 findings are public and 2,383 remain under embargo** — which is itself the story. A single model run has produced a backlog of undisclosed vulnerabilities larger than most national CERTs handle in a year, and the maintainers of those 269 projects now hold the timetable. GLM-5.3 was not downloadable at launch. Z.ai promised weights about two weeks after the announcement and lists the API as coming soon; for now the model runs through the GLM Coding Plan subscription and the ZCode agent. When the weights do land, the same capability that filled that ledger becomes available to anyone with the hardware to run it — which is the argument for staged release, and the argument against it, depending on who is making it. *Editorial note: wujec.ai reports on security capability as a published property of these models. We do not reproduce exploit material and we link only to the vendor's own coordinated-disclosure record.*

GLM-5.3 →