Skip to main content
Back to timeline
AnthropicSource publication:

Anthropic red team finds GLM-5.3 autonomously builds end-to-end exploits and its safeguards are bypassed 64%–100% of the time

Synopsis

Anthropic's Frontier Red Team evaluated GLM-5.3, the latest model from Zhipu AI (known outside China as Z.ai), using automated benchmarks and human-in-the-loop workflows, finding that it can autonomously build end-to-end cyber exploits (50 of 410 attempts on ExploitBench and full control-flow hijacks in 4% of trials on an internal binary exploitation benchmark) and that simple techniques bypassed its safeguards in 64%–100% of simulated tests, while those techniques did not succeed against safeguarded Claude models in their testing.

AI-generated editorial illustration: GLM-5.3 and the spread of advanced cyber capabilities

Interpretation

GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits, close to the level of Claude Mythos Preview. That capability had previously been demonstrated only by Claude Mythos Preview, released in a limited way through Project Glasswing; this work pairs a comparable level of exploit development capability with a publicly downloadable open-weight model. On ExploitBench, which measures exploitation of known vulnerabilities in the V8 engine used by Google Chrome, GLM-5.3 developed end-to-end exploits in 50 of 410 attempts versus 56 of 410 for Claude Mythos Preview; on 100 randomly selected tasks from the internal Binary Exploitation benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of trials versus 6% for Claude Mythos Preview, while earlier models such as Claude Opus 4.6 and GLM-5.2 did not succeed in any of them.

In human-in-the-loop offensive sessions, GLM-5.3 found and chained previously unknown vulnerabilities with very little human attention. Targets were chosen so the human experts were unaware of existing vulnerabilities, testing open-ended discovery and exploitation of novel flaws rather than reproduction of known issues. In one session, a researcher used GLM-5.3 on a sandboxed machine with a local Linux build of a popular web browser; over the course of a day and with limited human attention, the model found several previously unknown vulnerabilities in the browser's JavaScript engine and chained them into a working exploit: a webpage that, when visited, reads arbitrary files from the visitor's computer. The same session also identified exploitable vulnerabilities in wireless and graphics drivers and network-facing device software.

GLM-5.3's safeguards can be bypassed or removed, and removal leaves capabilities largely intact. Beyond reporting capability levels, the work quantifies how often safeguards are bypassed and reproduces the abliteration path specific to open-weight models. In simulated tests, simple techniques bypassed GLM-5.3's safeguards between 64% and 100% of the time, while none of these techniques got safeguarded Claude models to carry out the harmful tasks tested. Abliterating GLM-5.3 took the team about 2,200 GPU hours at roughly $4,400 in computation, and GLM-5.3-Flash about 600 GPU hours; the edit took the refusal rate from above 90% to about 3% on JailbreakBench, 2% on HarmBench, and 12% on StrongREJECT, while the standard and abliterated models scored the same on GPQA-Diamond and the abliterated version scored a few percent lower on a tested subset of CyberGym.

Turning a public fix into a working attack is very cheap. It shows how far N-day exploitation can be automated: the model converted public disclosure details into a working exploit chain with no significant human direction. Given public details of a Google Chrome flaw (CVE-2026-11645) and another known flaw, GLM-5.3-Flash chained exploits for the two flaws into a reliable exploit chain for an ARM64 target, bypassing pointer-authentication (PAC) hardening, taking 20 minutes of human attention plus 8 hours of model work, which at Zhipu's API prices would have cost $20.40.

Perspective

This work speaks to cyber defenders, model developers, and policymakers: it argues that with open-weight models now capable of end-to-end exploit development, defenders need frontier models at least as good as those their adversaries use, and the authors note efforts such as Project Glasswing and Patch the Planet to expand trusted defenders' access. The evaluations themselves run in isolated sandboxes against pre-set offline targets, so the findings characterize capability under those controlled settings rather than documenting real-world attacks.

Readers should still watch several things: the bypass rates come from simulated tests, and the authors note these techniques did not get safeguarded Claude models to carry out the harmful tasks tested, so behavior may differ across models and deployments; post-abliteration capability changes were measured only on GPQA-Diamond and a tested subset of CyberGym; the authors believe the browser vulnerabilities could also affect users on other platforms but say the path to exploitation there may be more complex and remains to be confirmed; and several vulnerability reports mentioned in the post are still under review and have not all been disclosed to maintainers.

Sources