The picture changes on ExploitBench, which requires deeper reasoning about how a real vulnerability could actually be exploited and demands a working exploit. GLM-5.3 more than doubles its predecessor's score, climbing from 24.4% to 54.4% . However, Z.ai reports that Mythos 5 scores 78% on ExploitBench and GPT-5.6 Sol scores 76.5% — leaving GLM-5.3 well behind on this more difficult measure
.
On ExploitGym, which counts how many exploitation tasks a model can complete under a normalized two-hour budget, GLM-5.3 finishes 105 tasks (up from 29 for GLM-5.2) . Within a six-hour budget, that rises to 130 tasks
.
| Benchmark | What It Measures | GLM-5.3 | Anthropic Mythos 5 | OpenAI GPT-5.6 Sol |
|---|---|---|---|---|
| CyberGym | White-box vulnerability discovery & validation from source code | 84.5% | 83.8% | 83.6% |
| ExploitBench | Root-cause reasoning + delivering a working exploit | 54.4% | 78% | 76.5% |
| ExploitGym (2-hour budget) | Exploit completions under normalized budget | 105 tasks | Not disclosed | Not disclosed |
Beyond benchmarks, GLM-5.3 has been deployed in real-world security testing with impressive results:
GLM-5.3 also posts strong numbers on coding and agentic benchmarks, positioning it as a leading open-weight model for software engineering tasks:
Z.ai says the model's cybersecurity capability "grew faster than anticipated" during post-training . The base model's 740 billion parameters were not retrained — all gains came from post-training improvements
. Because the capability was an emergent surprise, the company is holding back the open-weight release for two weeks while it strengthens safety and security controls
. The model is already available via the Z.ai API and through integrations with ZCode, Claude Code, and OpenCode
.