Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsGLM-5.3 performed close to Claude Mythos Preview on two exploit-development tests reported by Anthropic, but that is not proof the models have equivalent hacking ability overall. In a separate assessment, NIST’s Center for AI Standards and Innovation (CAISI) called GLM-5.3 the most cyber-capable open-weight model it had evaluated while estimating it lagged the U.S. frontier by about four months on its aggregate cyber measure. The results describe different tests, model sets and access conditions.
What Anthropic’s “Mythos-class” comparison means
Anthropic’s September 29, 2026 report evaluates GLM-5.3, an open-weight model developed by Zhipu AI, also known as Z.ai. Anthropic assessed that the model was released without meaningful safeguards against misuse. That characterization and the capability claims below are Anthropic’s findings, not a universal classification of the model.
“Mythos-class” is best understood as a shorthand for near results on two specific exploit-development evaluations. It does not establish broad equivalence with Claude Mythos Preview, nor does it show how often either model would succeed against real systems. The underlying tests measured particular outcomes in controlled environments.
How close were the models on exploit tests?
| Evaluation | GLM-5.3 | Claude Mythos Preview | What the result measures |
|---|---|---|---|
| ExploitBench, Anthropic’s reported run | 50 successful end-to-end exploits in 410 attempts | 56 successful end-to-end exploits in 410 attempts | Completion of exploits in this benchmark run—not all cybersecurity work or success against live targets. |
| Anthropic’s internal Binary Exploitation benchmark | Full control-flow hijacks in 4% of trials | Full control-flow hijacks in 6% of trials | A specific exploit outcome; it is not interchangeable with partial progress or vulnerability discovery. |
Anthropic also says GLM-5.2 and Claude Opus 4.6 did not succeed on these selected evaluations. That comparison applies to the evaluations Anthropic reported, not to every cyber task or model capability. Anthropic notes that some Claude models used for capability comparisons had safeguards disabled. Read Anthropic’s report and its evaluation details.
Recommended Free Tools
#1 Best Overall
What NIST’s CAISI found
In an assessment published September 17, 2026, NIST’s CAISI called GLM-5.3 the most cyber-capable open-weight model it had evaluated. On its aggregate measure across CAISI cyber benchmarks, it estimated that GLM-5.3 lagged the U.S. frontier by about four months. This is a date-bound estimate of aggregate benchmark capability, not a prediction that the model will trail by the same interval on every task.
The two organizations’ findings are not contradictory: Anthropic compared outcomes on two exploit-development evaluations, while CAISI assessed results across four benchmarks covering vulnerability discovery and exploit development. CAISI’s U.S. comparison also included models released only to trusted users, and its methodology disabled cyber safeguards on U.S. models where applicable. The comparison therefore is not a direct measurement of how publicly accessible models behave under default safeguards. See CAISI’s assessment and methodology.
What Anthropic reported about real-world-style workflows
Beyond automated benchmarks, Anthropic describes human-in-the-loop experiments in isolated, sandboxed settings. In one, researchers used GLM-5.3 in a Linux browser environment to find and chain previously unknown vulnerabilities. Anthropic says the work produced a proof-of-concept page able to read arbitrary files in that test setup, and that the vulnerabilities were disclosed to the maintainer.
Anthropic also says it tested GLM-5.3-Flash against known flaws. It reports that the model built an exploit chain over eight hours of model work, with about 20 minutes of human attention. These are company-reported controlled experiments, not verified intrusions against ordinary users. Anthropic cautions that simulated environments do not fully represent real operational conditions.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
How weak were the tested safeguards?
Anthropic tested GLM-5.3 in simulated malicious-order scenarios and reported these engagement rates:
| Test condition | Engagement rate reported by Anthropic |
|---|---|
| Deceptive red-team cover story | 64% |
| Prefilling of reasoning tokens | 92% |
| Abliterated copy, with model weights altered to reduce refusals | 100% |
These percentages describe Anthropic’s controlled test conditions. They are not estimates of how often real attackers would persuade the model to assist them. The distinction matters: a model’s behavior can change with prompt framing or altered weights, and a simulated malicious request is not the same as an actual attack in the wild.
Rank #4
Anthropic says its abliterated test copy was created by modifying GLM-5.3 to reduce refusals. The company reported about 2,200 GPU hours and roughly $4,400 in compute for its team’s experiment; it said an experienced team might need closer to 600 GPU hours and $1,200. Those figures describe this reported effort, not a standard market price or a turnkey action for ordinary users. Anthropic says refusal rates fell substantially across three harmful-request benchmarks while general capability remained largely intact on the checks it reported. Anthropic explains its safeguard tests and ablation experiment.
Quick Recap
Best Value
What readers should conclude
- There is credible evidence of advanced capability in specific tests: Anthropic reported results near Mythos Preview on two exploit-development evaluations, and CAISI independently found GLM-5.3 unusually capable among open-weight models it had evaluated.
- The evidence does not establish general equivalence: Benchmark endpoints, comparison models, safeguards and access conditions differ.
- The safeguard concern is conditional but material: Anthropic’s simulated tests found that prompting strategies and altered weights could substantially change behavior.
- Real-world risk is not quantified by these experiments: The reported tasks took place in controlled or sandboxed environments, which cannot establish how frequently a model would produce successful attacks in live conditions.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

