Anthropic’s sandboxed ExploitBench test produced working exploits in 50 of 410 GLM 5.3 attempts, versus 56 for Claude Mythos Preview.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Anthropic’s September 2026 analysis find about Z.ai’s open-weight GLM-5.3 model’s ability to build cyber exploits compared with Cla. Article summary: Anthropic’s September 2026 analysis found that Z.ai’s downloadable GLM-5.3 could autonomously build working, end-to-end exploits at nearly the rate of Claude Mythos Preview on one test—a capability Anthropic had previous. Topic tags: general, general web, government. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake n
Anthropic’s September 2026 evaluation found that Z.ai’s open-weight GLM-5.3 could build working, end-to-end exploits at a rate close to Claude Mythos Preview on one benchmark. The same report found that simple methods often bypassed GLM-5.3’s safeguards in simulations. Those results raise questions about how widely powerful cyber capabilities should be available—but they are controlled-test findings, not evidence of successful attacks against live targets. 12
On Anthropic’s ExploitBench evaluation, GLM-5.3 produced working exploits in 50 of 410 attempts; Claude Mythos Preview did so in 56. That is about 12% versus 14% on this test. The benchmark assessed exploit development against known vulnerabilities in a controlled setting, so the close scores should not be read as proof that the models perform equally across cybersecurity tasks. 5
12
Anthropic described GLM-5.3’s ability to build end-to-end exploits as a substantial advance over earlier models. Its report also compared the systems on a separate binary-exploitation evaluation: GLM-5.3 achieved full control-flow hijacks in 4% of tasks, compared with 6% for Mythos Preview. These results point to strong performance on particular tests, not universal equivalence. 3
7
12
NIST’s Center for AI Standards and Innovation called GLM-5.3 the most cyber-capable open-weight model it had assessed, while placing it about four months behind current U.S. frontier models on an aggregate of its cyber benchmarks. 17
That assessment does not directly conflict with Anthropic’s near-Mythos result: NIST compared GLM-5.3 across a broader collection of benchmarks, while Anthropic’s headline comparison was on a specific test against Mythos Preview. Different tests and comparison groups can produce different, compatible conclusions. 12
17
Anthropic reported that simple techniques bypassed GLM-5.3’s safeguards in 64% to 100% of its simulated tests. These percentages describe the results of the tested prompts and methods, not the probability that an attacker would succeed in a real-world operation. 12
The report also raised a separate issue with open weights: users who can access a model’s weights may be able to change its refusal behavior rather than just prompt around it. A secondary account of the reported experiment put the estimated cost of removing safeguards at about $1,200 for an experienced team, while also describing a roughly $4,400 compute estimate for the work. These are estimates tied to different descriptions of the effort—not a fixed price or a measure of how often real attackers would do it. 12
25
In one reported exercise, a researcher used the smaller GLM-5.3-Flash to build an exploit chain for two already-disclosed vulnerabilities. The work involved about 20 minutes of human attention, eight hours of model work and $20.40 in API charges, according to the report. One flaw was a Chrome vulnerability; the researcher supplied public details of the known flaws. This shows that the model could help develop a chain against disclosed vulnerabilities in a test—not that a live victim was compromised. 6
The distinction matters: a model’s ability to produce an exploit under test conditions is a warning about potential misuse, but it is not the same as demonstrating an attack in the wild. The available reporting here does not establish that GLM-5.3 was used to breach a real-world target. 5
6
Anthropic’s concern is that broadly available exploit-building capabilities can be used by attackers as well as defenders. The company has also argued for giving defenders access to advanced models so they can find and fix vulnerabilities. 1
12
Open weights create a difficult trade-off: they can make model capabilities available to independent researchers and defenders, while also giving users more ability to run or modify a model outside its developer’s controls. The benchmark and safeguard results inform that debate, but they do not by themselves show whether restricting access would make cybersecurity safer overall. 4
12
The strongest conclusion is narrower: in Anthropic’s evaluation, GLM-5.3 approached Mythos Preview on one exploit benchmark, and its safeguards proved vulnerable to tested bypass methods. NIST’s wider assessment still placed it behind current U.S. frontier models, and the reported demonstrations were not live attacks. 12
17
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Anthropic’s sandboxed ExploitBench test produced working exploits in 50 of 410 GLM 5.3 attempts, versus 56 for Claude Mythos Preview.