Morning Edition · Saturday, August 29, 2026Published at 2:21 AM EDT · New York
The Chinese lab delayed the download for two weeks after the model scored 84.5% on the CyberGym vulnerability benchmark, a score it says beat Anthropic's Mythos 5.

Zhipu AI, which sells internationally as Z.ai, has published the weights of GLM-5.3 for download, according to the Russian-language technical channel AI ML Big Data, which reported that the model can now be run locally and fine-tuned instead of being reached only through the company's hosted service. The release follows through on a delay the company set for itself. GLM-5.3 launched on August 14 with hosted access only, and Z.ai said at the time that downloadable weights would follow in roughly two weeks, once a safety review finished.
The reason for the pause was the model's performance on offensive security tasks. Z.ai reported a score of 84.5% on CyberGym, a benchmark that tests the ability to find and exploit software vulnerabilities, and said that result placed GLM-5.3 ahead of Anthropic's Mythos 5. The company also said its post-training run produced exploit-chain reasoning it had not intended to build, and that the model found more than a thousand critical bugs across Linux, WebKit and FreeBSD. Those counts come from Z.ai and have not been independently verified.
On coding, the improvement over the previous release is large by the lab's own accounting. GLM-5.3 keeps the same 744-billion-parameter base as GLM-5.2 and adds only post-training, more task environments and longer training runs, yet Terminal-Bench 3.0 rises from 4.6 to 28.3, with 28.5% on the command-line track of Agents' Last Exam. Z.ai calls that the state of the art among open-source models. Two days before releasing the flagship weights, the company released GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with 18 billion active parameters, under the MIT license at $0.15 per million input tokens and $0.50 per million output tokens.
The sequence matters more than any single number. A lab measured a dangerous capability, delayed publication, strengthened the model's defenses, and then released the weights anyway. Once weights are downloadable, the safety work becomes a starting point rather than a lasting control, because anyone with the right hardware can fine-tune it away.
Part of a tracked trend
Chinese Open-Weight Models Emerge as the Non-US AI Stack
As Washington restricts foreign access to US frontier models, governments and enterprises cut off from American AI increasingly standardize on downloadable Chinese open-weight models, splitting the world into competing AI supply blocs rather than a single frontier.
Start a discussion in Townsquare.
More from this edition
Z.ai gains standing as the open-weight lab that self-regulates and still ships, while enterprises and cloud providers that host models themselves gain leverage against the per-token pricing of OpenAI and Anthropic, and Nvidia gains demand for the hardware needed to run a 744-billion-parameter model locally.
The delay and the August 28 weight drop are corroborated by Axios and by the zai-org/GLM-5.3 repository, but every capability number, including the 84.5% CyberGym score and the 1,097 critical bugs, comes from Z.ai's own evaluation harness with no independent reproduction, and a safety delay that ends in publication also functions as marketing for the capability it warns about.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Open-weight releases now carry offensive security capability close to the frontier, and the decision to release sits with a lab outside United States jurisdiction, so American export policy cannot control it. Closed vendors selling coding agents lose pricing power against a model buyers can host themselves, and enterprise security teams gain both a defensive tool and a new threat model in the same download. The pattern of delaying and then publishing anyway also sets a template other labs will be measured against.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · Interconnects · Implicator
Comments
0No comments yet.