developing
Cybersecurity
Updated Sep 6, 2026, 12:24 PM UTC
Z.ai doubles exploit-analysis scores, delays weights for GLM-5.3
GLM-5.3 posts large gains on cyber benchmarks and a public vulnerability disclosure ledger, but Z.ai is withholding open weights until vetted security partners finish evaluation.
How this coverage works
This article combines reporting from 1 supporting Intel source. It is updated as material evidence arrives; prior published revisions remain in the record.
Z.ai has detailed GLM-5.3, which the company describes as its most capable model to date for cybersecurity tasks, alongside a staged release plan that withholds the model's full weights until after evaluation by vetted security partners. The announcement, posted to X, frames the release around a shift the company says is already underway: AI becoming part of both cyber offense and cyber defense, following an earlier incident in which GLM-5.2 helped Hugging Face investigate an AI system that had bypassed its own safeguards.
Z.ai reported substantial gains over GLM-5.2 across three benchmarks. On CyberGym, which starts from white-box source code and tests whether a model can identify and validate vulnerabilities by triggering faults, GLM-5.3 scored 84.5% versus 77.2% for its predecessor. On ExploitBench, which requires deeper reasoning about real vulnerabilities and how to exploit them, GLM-5.3 more than doubled GLM-5.2's score, reaching 54.4% against 24.4%. On ExploitGym, which measures completed exploitation tasks under a fixed evaluation budget, GLM-5.3 completed 105 tasks within two hours and 130 within six, compared with 29 and 39 for GLM-5.2.
The company says the GLM model series has already produced 2,436 vulnerability findings across 269 real-world projects in authorized evaluations conducted with universities and professional security teams, including 1,097 findings categorized as medium-to-high severity, spanning operating systems, browser engines, and network protocols. To manage disclosure of that work, Z.ai launched a "Security Disclosure Ledger" that records findings as they move through coordinated disclosure, publishing a cryptographic hash for issues still under embargo so a finding can later be verified without exposing operational details prematurely.
Z.ai says it will not publish GLM-5.3's model weights until staged evaluation with selected security partners and broader API access are complete. For safety, the company describes a three-layer approach: an external classifier that flags high-risk requests in its hosted services, a reasoning monitor that watches for harmful objectives emerging across multi-step tasks, and alignment built into the model itself so that safeguards travel with the weights into any local deployment. Z.ai is also launching an initiative called OpenVuln, through which it says it will work with maintainers of open-source projects to audit code, identify vulnerabilities, and support coordinated disclosure and remediation, aimed at smaller open-source projects that lack dedicated security resources of their own.
Update history1 updates
New facts extend one article instead of spawning duplicate write-ups across its source Intel pages.
Sep 6, 2026, 12:24 PM UTC
Revision 1 · initial
Initial published version covering the GLM-5.3 announcement, benchmark gains, and staged release plan.