Zhipu’s GLM-5.3 Claims Top Bug-Finding Scores Against GPT-5.6 and Claude
Chinese AI firm Zhipu says its GLM-5.3 model tops the CyberGym benchmark, finding 2,436 bugs across 269 real projects. Here's what the numbers actually show.

Chinese AI company Zhipu launched a new model called GLM-5.3 on August 11, 2026, claiming it beats competing models from OpenAI and Anthropic on CyberGym, a benchmark designed to test AI on real-world cybersecurity tasks. According to Zhipu, the model found 2,436 vulnerabilities across 269 real-world codebases, with 1,097 rated medium-to-high severity. The results raise serious questions about how quickly Chinese labs are closing the gap on western AI in security-sensitive domains, even while GLM-5.3 trails on other benchmarks.
What happened
| Detail | Figure |
|---|---|
| Model name | GLM-5.3 |
| Benchmark claimed | CyberGym (state of the art, per Zhipu) |
| Models it claims to beat | Fable 5 and GPT-5.6 Sol |
| Vulnerabilities found in real projects | 2,436 across 269 projects |
| Medium-to-high severity issues | 1,097 |
| Age of oldest bug found | Roughly 40 years |
| Areas covered | System kernels, OSes, browser engines, web apps, network protocols |
Zhipu, a Beijing-based AI lab, released GLM-5.3 with a specific focus on cybersecurity capability. CyberGym is a benchmark built to evaluate how well a model can solve real-world security challenges, not just textbook problems. Zhipu says GLM-5.3 tops that leaderboard above Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol.
In its announcement, Zhipu said the model’s gains were most pronounced further up what it calls the “exploitation chain.” Rather than flagging individual flaws in isolation, the model reportedly reasons across multiple attack stages and forms plans for complete exploitation sequences. That is a meaningful distinction: chaining vulnerabilities is how real-world attackers operate.
The company tested GLM-5.3 against actual codebases from Chinese companies. The bugs found span everything from operating system kernels to browser engines and network protocols. The oldest vulnerability had reportedly been sitting undetected for around four decades.
Where GLM-5.3 falls short
Zhipu’s own data shows GLM-5.3 does not lead across the board. The model scored below western competitors on other security and coding benchmarks. So this is a targeted performance claim on one specific test, not a general superiority argument. That nuance matters a lot when evaluating headlines that say “beats OpenAI and Anthropic.”
The CyberGym benchmark itself is worth scrutinising. Benchmarks in AI security are relatively new, and the community is still debating which ones reflect real-world attack and defence tasks most accurately. A top score on one benchmark, especially one where the model’s developer published the claim, warrants independent verification.
Why it matters
The speed of development is the real headline. Zhipu noted that “cyber capability developed faster than we expected” as it scaled post-training. The model’s debut closely follows Anthropic’s Mythos release, which means Chinese labs are reacting and iterating on western model capabilities within a very short window.
For businesses running software infrastructure, the practical implication is straightforward: AI-assisted vulnerability discovery is getting faster and cheaper, and it is no longer the exclusive domain of US labs. Whether GLM-5.3 is used offensively or defensively, the same underlying capability applies to both sides of that equation. Security teams that are not already using AI-assisted code scanning are falling behind the tools available to attackers.
This connects to a broader pattern we cover regularly in our AI news coverage: the gap between frontier western models and Chinese equivalents is narrowing faster than most analysts expected, particularly in specialised domains like security and code.
Our take
Take the benchmark claim with appropriate scepticism. Zhipu published this data itself, the CyberGym leaderboard is not widely audited, and the head-to-head comparison is narrow. That said, the real-world testing on 269 actual codebases is harder to dismiss. Finding 2,436 vulnerabilities, including bugs that survived undetected for decades, in genuine production code is concrete evidence of capability, not just benchmark theatre.
For any business running web applications or open-source infrastructure, this is a useful reminder that AI can now scan your code at a scale and speed that human security teams cannot match. If you are not already using automated vulnerability scanning as part of your build pipeline, that gap is worth closing now. If you want to explore how AI integration can strengthen your development workflow, that conversation is worth having sooner rather than later.
The bigger geopolitical point: any assumption that US AI dominance in security-relevant domains would hold for years looks shaky. GLM-5.3 is one data point, but the trajectory is consistent with what we have seen from Alibaba’s Qwen and other Chinese labs, as detailed in our coverage of Qwen’s rapid adoption growth.
What to do about it
- Add an AI-assisted static analysis tool to your CI/CD pipeline if you have not already.
- Run a one-time audit of your oldest codebases. Bugs that have lived there for years are exactly what these models find best.
- Watch the CyberGym leaderboard for independent third-party scores, not just vendor announcements.
- Brief your security team on exploitation-chain reasoning: the risk is not just isolated flaw detection but AI systems that plan multi-step attacks.
The clearest takeaway: AI vulnerability discovery is now a commodity capability, and your security posture should reflect that.
Frequently asked questions
What is Zhipu's GLM-5.3 model?
GLM-5.3 is an AI model released by Chinese company Zhipu in August 2026, focused on cybersecurity tasks. The company claims it leads the CyberGym benchmark for vulnerability discovery, outscoring models from OpenAI and Anthropic on that specific test.
What is the CyberGym benchmark?
CyberGym is a benchmark designed to evaluate how well AI models can solve real-world cybersecurity challenges, including vulnerability discovery and exploitation reasoning. It is distinct from general coding benchmarks.
How many vulnerabilities did GLM-5.3 find in real-world testing?
Zhipu says GLM-5.3 found 2,436 vulnerabilities across 269 real-world codebases, including 1,097 rated medium-to-high severity. The oldest bug discovered had reportedly gone undetected for roughly 40 years.
Does GLM-5.3 beat OpenAI and Anthropic models across all benchmarks?
No. Zhipu's own data shows GLM-5.3 scored below western models on other security and coding benchmarks. The top ranking is specific to CyberGym and should not be read as a general claim of superiority.


