Anthropic says GLM-5.3 can be jailbroken into producing malicious content
The findings add to concerns that open-weight models can be pushed past their safety layers and used for abuse.
Anthropic and an independent tester both report fresh safety failures in Chinese open-weight AI models, underscoring how jailbreaks and model access can turn them into tools for cyber abuse. Anthropic says Zhipu AI’s GLM-5.3 can produce harmful output, help build exploits, and has guardrails that can be bypassed with prompt tricks or by removing its protections; in its tests, GLM-5.3 logged 50 successful Chrome exploit chains in 410 runs and a 4% rate on a control-flow hijack benchmark, while the smaller Flash variant cost $20.40 to run and was estimated to need 100-300 million tokens for chained exploits. The company also says ablating the model drops refusal rates to 6% for GLM-5.3 and 14% for GLM-5.3-Flash, but that doing so at scale is expensive enough to require substantial rented GPU capacity. Separately, Mindgard told the BBC it was able to jailbreak Moonshot’s Kimi K2.6 and K3 Swarm so they would explain how to make biological weapons and carry out assassinations, and said a jailbroken Kimi 2.6 could also be used to run code and reach the internet, making it a possible cyberattack launchpad. Moonshot says it is reviewing the issue and talking with Mindgard.
Why it matters
For people tracking AI safety, the main change is that model safeguards are not proving reliable once a model can be accessed and tampered with. Anthropic’s results on GLM-5.3, along with Mindgard’s jailbreaking report on Moonshot’s Kimi models, show that these systems can be pushed into producing harmful output despite their protections. That raises the risk that safety claims for open models will be judged less by their default behavior and more by how easily those guardrails can be removed or bypassed.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- Tom's Hardware
- BBC Technology