GLM 5.3 exploit capabilities. The Chinese model GLM 5.3, developed by Z.ai, has demonstrated offensive cybersecurity capabilities comparable to those of Claude Mythos. This assessment was conducted by Anthropic, which published an analysis dedicated to the ability of language models to identify vulnerabilities and transform them into working exploits.

The analysis follows an independent evaluation by the National Institute of Standards and Technology (NIST) and examines both the technical performance of the models and the robustness of their integrated protections. A key element that emerged is the public availability of GLM 5.3’s weights, which allows for direct intervention on the model’s behavior, modifying the mechanisms designed to block potentially harmful requests.
GLM 5.3 exploit capabilities: why it matters
GLM 5.3 capabilities in vulnerability exploitation
To ensure test security and prevent impacts on real systems, Anthropic conducted all evaluations in isolated and protected environments (sandboxes). Using ExploitBench, a benchmark specifically designed for the development of exploits based on known vulnerabilities in Google Chrome’s V8 engine, GLM 5.3 was able to generate complete and functional exploits in 50 attempts out of a total of 410.
In an internal test focused on binary exploitation, the model achieved full control of the execution flow in 4% of cases over 100 attempts. These results demonstrate a concrete ability to move from defect analysis to the construction of a functional attack, although this does not automatically imply the possibility of compromising real systems.
Test sessions with expert researchers provided further qualitative details. Working on a Linux build of a web browser, the Chinese model identified several previously unknown vulnerabilities in the JavaScript engine and chained them into an attack capable of reading arbitrary files from the simulated victim’s computer. In the same context, exploitable flaws were also found in wireless drivers, graphics drivers, and software intended for network-exposed devices.
In an additional test, the GLM 5.3 Flash variant, starting from a known Chrome vulnerability, built an exploit chain for an ARM64-based system, successfully bypassing the Pointer Authentication protection mechanism.
NIST had already examined GLM 5.3 on September 17, 2026, defining it as the open-weight model with the greatest capabilities in the cyber domain among those analyzed by the agency. The evaluation was based on four specific benchmarks and a composite index constructed using an IRT (Item Response Theory) statistical model, which separates the difficulty of individual tasks from the intrinsic skills of the tested systems.
On SEC Bench Pro, the model correctly solved 40.4% of the proposed tests. On ExploitGym and OSS Fuzz, the results were 9.4% and 7.7%, respectively. According to NIST estimates, the performance gap compared to the most advanced US models in the aggregated index would be approximately four months.
Modifiable protections
A distinctive aspect of GLM 5.3 is the possibility of intervening directly on its internal weights. Researchers experimented with a technique called “abliteration,” achieving a significant reduction in refusals (refusal to respond to potentially harmful requests) without substantially compromising the model’s general performance.
The operation required approximately 2,200 hours of GPU computation and an estimated cost of about €3,800 in compute tokens, but for an expert team, the estimate reduces to approximately 600 GPU hours. After applying the technique, refusals dropped to 3% on JailbreakBench and 2% on HarmBench.
Even leaving the model’s weights intact, tests demonstrated the possibility of bypassing integrated protections. In a simulation, a request presented as “red teaming” activity (attack simulation to identify vulnerabilities) prompted the model to proceed in 64% of cases, with the percentage rising to 92% when accompanied by pre-filled reasoning.
Anthropic clarifies that the code generated during tests was not executed and that the simulated environment did not allow connections to external systems. This limits the scope of the conclusions but still highlights how easily the model’s behavior can be modified.
A protection system implemented via a remote service can be updated centrally, while a locally installed model offers greater control to the user employing it. Z.ai describes GLM 5.3 as an evolution obtained through post-training of the previous version, with a 50% improvement in its internal benchmark, Z.ai Code Bench.
For this reason, according to Anthropic, it is fundamental to evaluate not only a model’s capabilities at release but also the ease with which its operational constraints can be modified and bypassed.
Source and further reading on GLM 5.3 exploit capabilities: original article.
* Content created with the assistance of artificial intelligence systems.
Hardware Ready Ready to Bench?