Press Esc to close

Anthropic warns GLM-5.3 brings advanced cyber skills to open models

Anthropic's redacted GLM-5.3 browser exploit demonstration

Anthropic says Z.ai’s GLM-5.3 can build working cyber exploits at a level approaching the company’s Claude Mythos Preview in one benchmark. Its concern is that GLM-5.3 is available as an open-weight model, while its safeguards can be bypassed with relatively simple techniques.

In Anthropic’s ExploitBench evaluation of Chrome’s V8 engine, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview succeeded in 56. Anthropic says earlier models in its comparison did not complete an exploit in the benchmark. The company ran its evaluations in isolated environments, and says researchers disclosed newly found flaws to maintainers.

Anthropic also tested simulated malicious requests. GLM-5.3 engaged in 64 percent of trials when given a deceptive red-team pretext and 92 percent when its reasoning was prefilled. A modified version with refusals removed engaged in every trial. These are results from Anthropic’s controlled simulations, not estimates of real-world attack rates.

A separate assessment by the U.S. National Institute of Standards and Technology found GLM-5.3 was the strongest open-weight model it had evaluated for cyber tasks, while trailing the U.S. frontier by about four months on its aggregate benchmark. That offers useful context for Anthropic’s warning about public access to fast-improving capabilities.

Anthropic argues that governments should independently test highly capable models and that defenders need access to strong tools too. The debate over safeguards is also shaping how companies secure AI systems, as we reported when NVIDIA introduced its open agent safety platform.

Read Anthropic’s full GLM-5.3 assessment and NIST CAISI’s evaluation.

Comments