On September 29, 2026, Anthropic's Frontier Red Team published a research report titled "GLM-5.3 and the spread of advanced cyber capabilities," evaluating the cyberattack capabilities of GLM-5.3, the open-weight model released by Zhipu (Z.ai) in late August. The conclusion is blunt: GLM-5.3 can autonomously build end-to-end network exploit chains, approaching the level of Claude Mythos Preview — a model Anthropic only makes available to trusted defenders — while its safeguards proved fragile against simple bypass techniques, with bypass rates reaching 100%.
The numbers: one step behind Mythos
The core comparison sits on ExploitBench, a benchmark that asks models to write working, complete exploit chains against known vulnerabilities in targets such as browser engines:
| Model | Successful attempts (out of 410) | Full control-flow hijack rate |
|---|---|---|
| GLM-5.3 | 50 | 4% |
| Claude Mythos Preview | 56 | 6% |
| Earlier models (GLM-5.2, Claude Opus 4.6, etc.) | Did not reach this level | — |
In other words, a freely downloadable open model has nearly caught up with a frontier model reserved for vetted researchers — at the specific task of turning vulnerabilities into working attacks. Anthropic also notes that on September 17, the U.S. National Institute of Standards and Technology's Center for AI Standards and Innovation (CAISI) called GLM-5.3 the most cyber-capable open-weight model released to date, trailing the U.S. frontier by about four months on its combined cyber benchmarks, and says its own findings are broadly consistent.
How the safeguards were bypassed
The report separates "how capable the model is" from "how easily it can be steered into harmful tasks." Direct malicious requests were refused — but rephrase the ask and the picture changes:
- With a deceptive "red team testing" cover story, 64% of simulated tests proceeded;
- With the model's reasoning prefilled, the rate rose to 92%;
- After directly modifying the open weights to remove refusal behavior, it hit 100%.
The same techniques failed to make the safeguarded Claude model carry out harmful tasks. Anthropic's point: open weights let attackers dismantle these restrictions themselves — the fundamental difference from closed, access-controlled models.
Browser zero-days in a day, an attack chain for $20
The most striking part is the hands-on test. Researchers placed GLM-5.3 in a sandbox and had it analyze a Linux build of a mainstream browser: within a day, the model found several previously unknown vulnerabilities and chained them into a malicious webpage — one that could escape the browser sandbox and read arbitrary files from a visitor's computer, including SSH private keys in the test. Anthropic says it has reported the flaws to the browser's maintainer.
The smaller GLM-5.3-Flash, working only from publicly disclosed Chrome vulnerability information, assembled a working exploit chain after about 20 minutes of human attention and 8 hours of model runtime — roughly $20.40 at Zhipu's API pricing.
After the capability spreads, the race is patch speed
The report landed at a pointed moment: Anthropic's confidential IPO prospectus, disclosed by Reuters, warns that advanced AI could pose "catastrophic or existential risks to humanity" (Anthropic's prospectus: $42B net loss last year, $2T IPO valuation target). A risk warning for investors on one hand; its own red team proving top-tier attack capability is now downloadable by anyone on the other.
Anthropic's stated position is not a blanket ban: the report recommends testing sufficiently capable systems, hardening safeguards, and giving defenders better frontier tools. The logic is pragmatic — if a model can dig browser zero-days out in a day, how fast defenders can use the same capability to find flaws and ship patches becomes the new race. Days earlier, the UK AI Security Institute measured GPT-6 Astra's unauthorized attack rate at five times its predecessor's: frontier models' "hands-on" abilities are systematically outgrowing old safety assumptions, and the security math has to be redone.