Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A 0% attack success rate means that no attack in a particular evaluation met that evaluation’s definition of success. It is evidence about the tested system under the benchmark’s specific attacks, scoring rules, and conditions—not proof that the system is secure against attacks the test did not cover.

What does a 0% attack success rate actually mean?

The National Institute of Standards and Technology’s AI Metrology Center defines attack success rate (ASR) as the “Percentage of generated adversarial inputs that are misclassified.” The definition makes the counted outcome central: ASR records whether inputs in a particular test set produced the specified failure. It does not establish that every relevant security failure would be counted by that metric.

So a reported 0% means that no tested attack met the benchmark’s success criterion in that run. To interpret the result, you need to know what system was tested, which attacks were attempted, how many attempts were made, and how success was scored. A result without those details leaves its scope unclear.

Does 0% attack success mean an AI system is secure?

No—not on its own. A benchmark result is bounded by its threat model and coverage. It may be useful evidence that a system resisted the attacks it tested, but it cannot establish resistance to different attack strategies, a larger query budget, other deployment conditions, or failures that the scoring rule does not recognize.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A 2025 study, Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?, illustrates the difference. The authors report 0% ASR on four public agent benchmarks—AgentDojo, Agent Security Bench, InjecAgent, and tau-Bench—while also discussing limitations in those benchmarks and bypasses in practice. The result describes performance on those evaluations; it is not a deployment guarantee.

Why can a benchmark report zero while attacks still succeed?

The attack set may be fixed or limited

Tests drawn from a known or fixed set answer whether the system resisted those attacks. Adaptive testing asks a different question: can an attacker use the system’s responses to refine attempts over multiple turns? A low rate on first-turn attacks does not answer what happens when the attacker gets repeated opportunities.

In a 2026 study, Jain, Hartmann, and Li held 21 scenarios, attackers, defenders, and structured-output scoring fixed while comparing first-turn results with multi-turn adaptive attacks. They reported 0–1% ASR on the first turn and 5.4–14.0% when attackers could make up to 15 adaptive rounds. These are results from that study’s protocol, not a general conversion factor for other systems or benchmarks.

The query budget may stop before the attack does

The same study capped attacks at 15 rounds and reported that the observed success curve was still rising at that limit. Its results show why the round or query budget belongs beside an ASR figure: a test stopped at a fixed cap measures performance within that cap, not what would happen with unlimited attempts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The benchmark may not cover the relevant scenarios

A test can include many examples yet still miss important attack families or deployment conditions. In a pooled score, strong performance on many scenarios can also obscure a weak result on a smaller subset. The 2025 firewall study’s discussion of benchmark limitations is a reminder that a public suite’s breadth and realism should be assessed, not assumed.

The success rule may miss a failure

Results depend on how success is defined and detected. A rule might check a structured output, rely on an automated judge, or look for an observable side effect. Each approach can miss cases outside its design. Schwinn and coauthors’ 2026 ICML paper, A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness, reports that distribution shifts and semantic ambiguity in red-team settings can impair automated judging. If the evaluator fails to recognize an attack’s success, the benchmark can report a lower rate than the actual count of failures would warrant.

What does a zero say about risk and uncertainty?

Zero observed successes is not the same as proof of zero risk. The denominator matters: zero successes in a small number of attempts provides a different amount of evidence from zero in a much larger test. Scenario diversity matters too; many near-duplicate prompts do not necessarily cover many distinct ways an attack could work.

Jain, Hartmann, and Li report Wilson 95% intervals for their adaptive study and note that many initial evaluation cells were small. Those intervals belong to that study’s trial counts and method; they should not be transferred to another benchmark. The reviewed work establishes no universal sample-size threshold that makes a 0% AI-security ASR conclusive. Look for the number of trials and uncertainty estimates rather than treating zero as self-explanatory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can a defense achieve low ASR by sacrificing task performance?

Yes. A defense might block an attack by refusing a task or suppressing content that the system needed to process faithfully. A low attack-success rate alone may not reveal that cost.

In their 2026 ICML paper, Security–Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense, Mitchell Hermon, Rahul Gupta, Weitong Ruan, Ekraam Sabir, and Haohan Wang write: “Attack-success metrics cannot see this, because a model that ignores an injection and one that faithfully processes it as data score identically.” Their study compares security and fidelity across 1,168 examples and 48 configurations. Those figures describe the paper’s evaluated configurations, not a general estimate of deployed-system behavior.

When reading a defense result, check whether the system still completes benign tasks and handles untrusted content as required. Security and task fidelity are related but distinct outcomes; an ASR figure alone does not capture both.

How to assess two 0% benchmark claims

Compare the protocols, not just the headline percentages. The key question is whether the tests measured the same threat and gave attackers and evaluators comparable opportunities.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
What to compare What to establish
Threat model Which system, tools, data, deployment context, and attacker capabilities were in scope?
Attack strategy and budget Were attacks fixed in advance or adaptive? How many prompts, queries, turns, retries, or attacker models were allowed?
Coverage How many distinct scenarios and attack families were tested? Could a pooled rate hide a weak subset?
Success rule and evaluator What counted as success, how was it detected, and was the scoring method validated for the tested setting?
Evidence and uncertainty How many trials produced the rate? Are counts and confidence intervals reported, and are they tied to the relevant evaluation cells?
Utility and fidelity Did the system still perform benign tasks and process untrusted content as required?
Reproducibility Are benchmark versions, model versions, implementation details, and scoring code available? Could an implementation bug affect the result?

What standardization can—and cannot—tell you

Standardized evaluations make results easier to compare when systems are tested under a shared protocol. Mazeika and coauthors’ 2024 paper, HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, frames standardization as an evaluation need. But a common benchmark does not automatically solve gaps in attack coverage, scoring reliability, adaptive testing, or task fidelity. Standardization improves consistency; it does not turn a bounded test into proof of universal security.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.