# Z.ai Releases GLM-5.3 Weights After a Safety Hold on Its Cyber Capabilities

The 743-billion-parameter mixture-of-experts model posts 28.3 on Terminal-Bench 3.0 and 84.5 on CyberGym, and Z.ai says it delayed the download because exploitation skills improved faster than expected.

- Published: 2026-08-30T06:24:05.945Z
- Canonical: https://polylog.news/ai/2026-08-30/z-ai-releases-glm-5-3-weights-after-a-safety-hold-on-its-cyb
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [Decrypt](https://decrypt.co/375684/china-z-ai-glm-5-3-top-open-weight-coding-model), [Interconnects](https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride), [Hugging Face](https://huggingface.co/zai-org/GLM-5.3)

Z.ai has [published the weights for GLM-5.3](https://twitter.com/Zai_org/status/2093354097122455713), the coding and agent model it launched through its application programming interface (API) and subscription coding plan earlier this month. The repository on [Hugging Face](https://huggingface.co/zai-org/GLM-5.3) is public and ungated, with an FP8 checkpoint of about 756 gigabytes and a BF16 release near 1.5 terabytes. A Russian-language technical channel [flagged the drop](https://t.me/ai_machinelearning_big_data/10799) as a full weights release rather than API access, meaning teams can run and fine-tune the model on their own hardware.

The model is a 743-billion-parameter mixture-of-experts design with roughly 40 billion parameters active per token, and it shares its base with GLM-5.2. Z.ai says every gain came from post-training: more environments, more task diversity, and more compute spent on reinforcement learning over that stack. The company reports 28.3 on Terminal-Bench 3.0, up from 4.6 for the prior release, and 66.9 on DeepSWE v1.1, and it calls the model [the strongest open-weights coder available](https://decrypt.co/375684/china-z-ai-glm-5-3-top-open-weight-coding-model). Those numbers come from the vendor, and independent reproduction on the long-horizon command-line benchmark has not yet been published.

The security profile is the more consequential part. GLM-5.3 scores 84.5 on CyberGym, a benchmark for finding software vulnerabilities, which places it marginally above Claude Mythos 5 and GPT-5.6 Sol on that measure, and Z.ai says its largest gains sit further along the exploitation chain. The company delayed the release after the API launch to run what it described as its most extensive risk review to date, because offensive-security capability had grown faster than it expected. Once weights are downloadable, that review governs nothing about how the model is used.

Analyst [Nathan Lambert argued](https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride) that Chinese labs are now matching the frontier on the practical agentic work engineers actually buy, rather than trailing by a generation. A same-base, post-training-only jump of this size supports that read, and it also suggests the remaining headroom in reinforcement learning environments is large enough that a lab without a new pretraining run can still move the frontier of open models.

## What this means

A downloadable model that matches closed frontier systems on vulnerability discovery removes the gating point that API vendors rely on. Closed labs price and police cyber capability through refusal training and usage policy at the endpoint, which no longer binds anyone who runs 756 gigabytes of weights locally. The immediate winners are enterprises and governments that want frontier-grade coding agents without a US vendor relationship. The exposed parties are the closed-model API businesses, whose premium pricing rests partly on capabilities customers cannot otherwise obtain, and every defender whose threat model assumed exploitation skill stayed behind a provider's safety filter.

## What to watch

- Whether an independent group reproduces the Terminal-Bench 3.0 and CyberGym numbers on the downloaded weights, which is what separates a genuine capability step from a favorable evaluation harness.
- The license terms attached to the Hugging Face repository, since Z.ai shipped the prior release under permissive terms and a narrower license would change what commercial users can legally deploy.
- Whether US or European regulators respond to an openly downloadable model that scores at frontier level on offensive-security benchmarks, which would be the first real test of applying export or release rules to weights rather than to services.
