Morning Edition · Sunday, August 30, 2026Published at 2:24 AM EDT · New York
The 743-billion-parameter mixture-of-experts model posts 28.3 on Terminal-Bench 3.0 and 84.5 on CyberGym, and Z.ai says it delayed the download because exploitation skills improved faster than expected.

Z.ai has published the weights for GLM-5.3, the coding and agent model it launched through its application programming interface (API) and subscription coding plan earlier this month. The repository on Hugging Face is public and ungated, with an FP8 checkpoint of about 756 gigabytes and a BF16 release near 1.5 terabytes. A Russian-language technical channel flagged the drop as a full weights release rather than API access, meaning teams can run and fine-tune the model on their own hardware.
The model is a 743-billion-parameter mixture-of-experts design with roughly 40 billion parameters active per token, and it shares its base with GLM-5.2. Z.ai says every gain came from post-training: more environments, more task diversity, and more compute spent on reinforcement learning over that stack. The company reports 28.3 on Terminal-Bench 3.0, up from 4.6 for the prior release, and 66.9 on DeepSWE v1.1, and it calls the model the strongest open-weights coder available. Those numbers come from the vendor, and independent reproduction on the long-horizon command-line benchmark has not yet been published.
The security profile is the more consequential part. GLM-5.3 scores 84.5 on CyberGym, a benchmark for finding software vulnerabilities, which places it marginally above Claude Mythos 5 and GPT-5.6 Sol on that measure, and Z.ai says its largest gains sit further along the exploitation chain. The company delayed the release after the API launch to run what it described as its most extensive risk review to date, because offensive-security capability had grown faster than it expected. Once weights are downloadable, that review governs nothing about how the model is used.
Analyst Nathan Lambert argued that Chinese labs are now matching the frontier on the practical agentic work engineers actually buy, rather than trailing by a generation. A same-base, post-training-only jump of this size supports that read, and it also suggests the remaining headroom in reinforcement learning environments is large enough that a lab without a new pretraining run can still move the frontier of open models.
Part of a tracked trend
Open-Weight Models Close the Gap With Closed Frontier Labs
Over the next 3-9 months, open-weight releases with downloadable weights, long context, and strong agentic/coding performance increasingly match closed frontier models on practical work, eroding the closed-lab moat.
Start a discussion in Townsquare.
More from this edition
Z.ai and China's open-weight labs gain standing as frontier-equivalent suppliers, enterprises and governments that want top-tier coding agents without a US vendor contract capture the saving, and the closed-API vendors lose part of the capability premium their pricing rests on.
Independent outlets corroborate the two-week hold and the 84.5 CyberGym score, which puts GLM-5.3 only about a point above Claude Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6, but every headline figure originates with Z.ai, and a safety delay the company announced publicly also served as an advertisement for the exploitation capability it said it was reviewing.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
A downloadable model that matches closed frontier systems on vulnerability discovery removes the gating point that API vendors rely on. Closed labs price and police cyber capability through refusal training and usage policy at the endpoint, which no longer binds anyone who runs 756 gigabytes of weights locally. The immediate winners are enterprises and governments that want frontier-grade coding agents without a US vendor relationship. The exposed parties are the closed-model API businesses, whose premium pricing rests partly on capabilities customers cannot otherwise obtain, and every defender whose threat model assumed exploitation skill stayed behind a provider's safety filter.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · Decrypt · Interconnects · Hugging Face
Comments
0No comments yet.