GLM-5.3 brings open-weight AI closer to cyber exploitation

GLM-5.3 emerges as a powerful open-weight cybersecurity AI after finding thousands of vulnerabilities

Z.ai Is Preparing to Release an AI Model With Serious Cyber Capabilities​

Chinese AI company Z.ai has unveiled GLM-5.3, a new coding model whose most consequential upgrade may be its ability to find and investigate software vulnerabilities. Z.ai describes the model's cybersecurity behavior as an emergent capability rather than the original purpose of the system, and says it has already used GLM-5.3 with security teams to identify thousands of vulnerabilities in real open-source projects. The company plans to make the model publicly available after an additional security-review period of about two weeks.

GLM-5.3 scored 84.5% on Z.ai's CyberGym evaluation​

According to Z.ai's GLM-5.3 launch report, the model reached 84.5% on its CyberGym evaluation, up from 77.2% for GLM-5.2. In the comparison published by Z.ai, that put GLM-5.3 narrowly ahead of Anthropic's Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.
The ranking needs an important qualification. These are Z.ai-reported results and have not been independently verified. CyberGym also measures a specific part of the security workflow, particularly a model's ability to inspect code, identify vulnerabilities and validate findings. A narrow lead on that benchmark does not establish GLM-5.3 as the strongest cybersecurity model across every task.
The more useful conclusion is that an open-weight candidate is now performing close to tightly controlled frontier systems on a serious vulnerability-research benchmark.


Its offensive capability is powerful, but not benchmark-leading​

The most revealing result comes from comparing vulnerability discovery with exploit development. Z.ai reports a 54.4% score for GLM-5.3 on ExploitBench, while Mythos 5 reached 78.0% and GPT-5.6 Sol reached 76.5%. That gap complicates claims that GLM-5.3 is already the world's most capable AI hacker.
At the same time, Z.ai says the model demonstrated much stronger long-horizon cyber behavior than its predecessor, completing 105 attack-development tasks in a two-hour evaluation and 130 over six hours. The company describes these abilities as emerging from broader improvements in coding, reinforcement learning and long-running agentic work rather than from building a dedicated offensive security model.
The distinction matters. GLM-5.3 appears exceptionally strong at finding vulnerabilities, while the harder task of turning those findings into reliable offensive outcomes remains an area where restricted frontier systems retain a substantial advantage.


The real-world audit found 2,436 vulnerabilities​

Benchmark scores are only part of Z.ai's case. The company has also worked with security teams to examine real software projects, and its public vulnerability disclosure ledger currently records 2,436 findings across 269 open-source projects.
Of those findings, 107 are classified as critical and 990 as high severity, producing a combined total of 1,097 critical or high-severity vulnerabilities. The registry currently shows 53 vulnerabilities as publicly disclosed while another 2,383 remain undisclosed, allowing affected projects time to address the issues before details become public.
The projects span major categories of software infrastructure rather than a narrow collection of experimental repositories. That makes the audit a more meaningful signal than benchmark performance alone because the model had to operate against large, mature codebases with years of accumulated complexity.


Some of the bugs had survived for decades​

The age of the findings may be even more striking than their number. Z.ai's disclosure database says the vulnerabilities span 45 years, with the oldest traced to code introduced in 1981. Across the full dataset, the average vulnerability had remained undiscovered for 26.6 years before being identified.
That is an important indicator of what advanced AI security agents could change. Traditional audits are constrained by human time, specialist availability and the cost of repeatedly reviewing enormous legacy codebases. An agent capable of sustaining repository-scale analysis can revisit software that has already survived years of human inspection and still find overlooked weaknesses.
The implication is not that human security researchers become unnecessary. It is that the amount of code a security team can realistically inspect may expand dramatically when capable agents become persistent collaborators rather than occasional assistants.


Z.ai is delaying the open-weight release because of cyber risk​

The unusual part of the launch is what Z.ai is not doing immediately. Instead of releasing the model weights alongside the announcement, the company says it plans to wait about two weeks while completing additional safety assessments and strengthening safeguards. Reuters reported that Z.ai also intends to use a trusted-access program for its most sensitive cybersecurity capabilities.
The delay illustrates the conflict at the center of the release. Open weights give security researchers, smaller companies and open-source maintainers access to capabilities that might otherwise remain concentrated inside a few closed AI labs. The same openness also makes centralized safety controls much harder to enforce once a model can be downloaded, modified and connected to external tools.
How Z.ai reconciles an open-weight release with restricted access to the most sensitive cyber workflows may become more important than another benchmark point.


Open Source Shield turns the model into a security strategy​

Z.ai is framing GLM-5.3 as more than a model release. The company has announced an Open Source Shield initiative intended to use its models to audit selected open-source projects, provide defensive access to security teams and expand automated code-review capabilities in its development products.
The argument is straightforward: if powerful cyber AI remains available only to a handful of major corporations and vetted institutions, maintainers of widely used open-source software may face advanced attackers without comparable defensive tooling. An open model could narrow that capability gap.
The opposite risk is equally obvious. Once capable weights circulate broadly, the same economics that make large-scale defensive auditing inexpensive can also reduce the cost of malicious vulnerability discovery. The technology does not inherently know which side of that competition will use it more effectively.


GLM-5.3 pushes the AI security race into a harder phase​

The significance of GLM-5.3 is not that an autonomous AI hacker has suddenly become available to everyone. The model is not yet publicly released, its headline benchmark results come from Z.ai's own evaluation, and its exploit-development scores remain well below the strongest restricted systems in the comparison.
What has changed is the distance between open models and controlled cyber systems. A model scheduled for public release is already demonstrating strong vulnerability discovery, long-running agent behavior and real-world results across hundreds of open-source projects. That makes the traditional divide between general coding AI and specialized cybersecurity tooling much less clear.
The next two weeks will therefore matter for more than Z.ai. Once the weights leave the company's infrastructure, the industry will get a much clearer test of whether advanced cyber capability can remain meaningfully governed after it becomes open.



Editorial Team - CoinBotLab
  • Reading time 6 min read
  • Views3
  • Reading time 6 min read
  • Views3
  • Reading time 5 min read
  • Views8
  • Reading time 5 min read
  • Views6
  • Reading time 5 min read
  • Views10
  • Reading time 5 min read
  • Views21

Comments

There are no comments to display

Information

Author
Coinbotlab
Published
Reading time
6 min read
Views
3

More by Coinbotlab

Top