Summary

Zhipu released GLM-5.3, built on the same base as GLM-5.2 with all gains coming from post-training ('scaling post-training is all we did'). It is positioned as the strongest open-weights coding model, with a heavy cybersecurity and agentic focus: working with Chinese security teams, it found 2,436 vulnerabilities across 269 projects, published in the public registry cvd.z.ai. The model is available now via the GLM Coding Plan, which works with Claude Code and OpenCode. The API is 'coming soon', and the open weights are planned for ~2 weeks out, after security reviews.

Why it matters
For coding-agent builders this is a useful data point: post-training alone is claimed to have taken open-weights coding SOTA and beaten US frontier models on CyberGym, suggesting post-training scaling can still move agentic and coding capability substantially. The pattern of releasing weights only after a security review is itself worth watching.
Technical details
Context 1M tokens, 128K max output
Zai Code Bench +50% vs GLM-5.2
Terminal Bench 3 0 28.3 (from 4.6)
Deepswe V1 1 66.9 (from 46.2)
Agents Last Exam 28.5 (from 23.8)
Cybergym 84.5% (vs Mythos 5 83.8%, GPT-5.6 Sol 83.6%)
Exploitbench 54.4% (from 24.4%)
Vulnerabilities Found 2,436 across 269 projects (registry: cvd.z.ai)
Weights landed 2026-08-25: zai-org/GLM-5.3 (FP8, 153 files) + GLM-5.3-BF16, repos created 2026-08-25T06:42Z (~3 days ahead of the ~8/28 estimate); ungated; 753B total params
License custom 'glm-5.3' license: MIT-style; single restriction — Model-as-a-Service businesses with >$10B annual revenue must pass Z.ai security review before commercial use
Independent Eval Artificial Analysis Intelligence Index 60 (GLM-5.3 max; on par with Kimi K3, +7 vs GLM-5.2); GLM-5.3-Flash 57 (added 2026-08-26)
Updates
2026-08-15 Z.ai published follow-up 'Preparing GLM-5.3 for Open Release: A Responsible Path to Cyber Defense' at 2026-08-15T06:32Z, plus live OpenVuln demo space (huggingface.co/spaces/zai-org/OpenVuln). Reaffirms weights-after-review plan.
2026-08-25 Recovered from index (previous run wrote the update note to index/events.json only): open weights landed: HF zai-org/GLM-5.3 (FP8, 153 files) + GLM-5.3-BF16 created 2026-08-25T06:42Z (~3 days ahead of the ~8/28 estimate); ungated; 753B total params; custom glm-5.3 license (MIT-style; only restriction: >$10B-revenue MaaS businesses need Z.ai security review before commercial use); independent eval: Artificial Analysis Intelligence Index 60 (max) — on par with Kimi K3, +7 vs GLM-5.2; captured 2026-08-28 run; recommendation stays TRIAL
2026-08-29 Model card on HF fully rewritten (repo 'Initial commit 0828' at 2026-08-27T17:16:16Z, coverage-gap recovery): complete benchmark table with per-benchmark footnotes is now public — states GLM-5.3 uses the same base model as GLM-5.2 (all gains from post-training); Terminal Bench 2.1 88.2 / TB 3.0 28.3 / DeepSWE v1.1 66.9 / CyberGym 84.5 / ExploitGym 105@2h / HLE w-tools 62.5 / GDPval-AA v2 1769; FrontierSWE 78.1 evaluated by Proximal; most evals run in a Claude Code 2.1.207 harness. Community PRs #2 (2026-08-28T15:22Z) and #3 (2026-08-29T09:51Z) added .eval_results YAML metadata mirroring model-card numbers (source field: 'Model Card') — HF metadata normalization, NOT independent reproduction of the weights; trend #1 criterion (a) remains open. Downloads @2026-08-29T11:35Z: FP8 repo 8,804 (30d, effectively all-time) + 1,207 likes; BF16 repo 482. Recommendation stays TRIAL.
2026-08-31 Adoption & reproduction check @2026-08-31T00:00Z: FP8 repo downloads jumped from 8,804 (8/29 sample) to 50,116 (~5.7x in ~2 days), likes 1,207 -> 1,336 — the 753B flagship artifact is now being pulled at scale alongside the Flash variant. HF discussions #7-#9 (8/29-30) remain minor (release-cadence question, citation error, praise); still no third-party Terminal Bench / SWE reproduction of the weights — trend #1 criterion (a) unmet. Recommendation stays TRIAL.
2026-09-02 adoption & reproduction check @2026-09-02T00:0xZ: FP8 repo 50,116 -> 94,403 downloads (+88% in ~2 days, ~8 days after listing), likes 1,336 -> 1,466; BF16 repo now also public. HF discussions #10-#15 (8/31-9/1) are still metadata-sync PRs (TB 2.1/3.0 harness notes, Toolathlon-Verified eval result) and refusal complaints ('Safety Maxxed to oblivion') — no third-party Terminal Bench / SWE reproduction; trend #1 criterion (a) still unmet. Ecosystem note: Baseten's 9/1 'efficient frontier of LLM inference' blog uses GLM-5.3 as its canonical agentic-coding serving example (conceptual framework; no Flash-specific throughput numbers). Stays TRIAL
2026-09-05 Adoption & reproduction check @2026-09-05: FP8 repo 94,403 -> 303,534 downloads (+222% in ~3 days, 30d window), likes 1,466 -> 1,695; unsloth/GLM-5.3-GGUF derivative at 164,424. Discussions #16-#19 (9/2-9/4): #16 merged eval-results metadata PR (HF staff, terminal-bench-3.0 dataset pointer), #17 spam upload, #18 an independent quantization-fidelity measurement (KL vs reference on frozen tokens) — a genuine independent measurement but NOT a Terminal-Bench/SWE reproduction, #19 instruction-following complaint; trend #1 criterion (a) remains unmet. Stays TRIAL
Tags
glmzhipucodingagentscybersecurityopen-weights-pending