Anthropic: Zhipu AI's GLM-5.3 Has Weak Safeguards, Mythos-Class Hacking Skills
Anthropic's frontier red teaming report claims Zhipu AI's open-weight GLM-5.3 can generate malicious content and be used for cyberattacks.
Anthropic published a frontier red teaming report claiming that Zhipu AI's open-weight GLM-5.3 model has weak safeguards and can be used for cyberattacks. The report says the model can generate malicious content and that its safeguards can be bypassed through deceptive prompts, prefilling thinking tokens, and abliteration.
In Anthropic's Exploitbench, GLM-5.3 developed end-to-end exploits 50 times in 410 runs, compared with 56 for Anthropic's Claude Mythos. In an internal benchmark for full control-flow hijacks, GLM-5.3 had a 4% success rate versus 6% for Mythos, while Kimi K3 and DeepSeek V4.1 Flash scored 0%.
Anthropic also says that abliterating GLM-5.3 dropped its refusal rate to 6%, and GLM-5.3-Flash to 14%. The report was covered by tomshardware.com.
Earlier warnings
In late September, the Center for AI Standards and Innovation published a report claiming GLM-5.3 can fully automate exploits at a level similar to Anthropic's unreleased Claude Mythos model.
Anthropic developed Project Glasswing to give developers access to a Mythos-class AI model to patch bugs and fix vulnerabilities before such models are released. The company also released Claude Opus 5.5 and Claude Sonnet 5.5 days after alarms were raised about the pace of AI development.
Quick answers
What does Anthropic's report say about GLM-5.3?
Anthropic claims GLM-5.3 can generate malicious content and be used for cyberattacks, with safeguards that can be bypassed via deceptive prompts, prefilling thinking tokens, and abliteration.
How did GLM-5.3 perform in Anthropic's Exploitbench?
GLM-5.3 developed end-to-end exploits 50 times in 410 runs, versus 56 for Claude Mythos.
What happened to GLM-5.3's refusal rate after abliteration?
Anthropic says abliterating GLM-5.3 dropped its refusal rate to 6%, and GLM-5.3-Flash to 14%.