| Benchmark | GLM-5.3 | GLM-5.2 (prior generation) |
|---|---|---|
| Terminal-Bench 3.0 | 28.3 | 4.6 |
| DeepSWE v1.1 | 66.9 | 46.2 |
| Agents' Last Exam (CLI) | 28.5 | 23.8 |
| Z.ai Internal Code Bench | ~50% improvement | baseline |
Note: Mythos 5 scores 80.3% on SWE-bench Pro (agentic coding), but direct comparisons on these specific benchmarks are not available across both models . All of GLM-5.3's gains came from expanded post-training only — the base 743B model was not retrained
.
GLM-5.3 claims a narrow lead over Mythos 5 on vulnerability detection (84.5% vs 83.8% on CyberGym), but trails significantly on exploitation (54.4% vs 78.0% on ExploitBench). Its open-weight release is delayed for safety review, and its predecessor GLM-5.2 was assessed as having near-zero refusal of cyberattack instructions. Mythos 5 remains far more restricted — never publicly available, limited to ~200 vetted partners — but also far more capable at the most dangerous offensive task: generating working exploits. The broader implication is that open-weight models are rapidly closing the capability gap with frontier restricted models on cyber tasks, while safety safeguards remain substantially weaker. Chinese regulatory uncertainty about exporting top AI models adds another layer of complexity to whether GLM-5.3's weights will ever be as freely available as Z.ai's marketing suggests.