Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a benchmark of 10 open-source prompt-injection detectors against 629 attacks hidden in ordinary AI-agent tool output, the best default balance belonged to jailbreak-detector-large: it caught 319 attacks (51%) and flagged 2 of 97 benign outputs (2%). The result measures text classification—not whether a live agent would obey an attack or whether a dangerous action would be stopped.

What the benchmark tested

Rudratosh Shastri’s buried-injections benchmark, reviewed October 5, 2026, evaluates whether detectors identify attack text embedded in otherwise ordinary tool output. Its default leaderboard covers 629 attack examples and 97 benign examples. It also tests the 27 distinct attack texts on their own, without the surrounding tool-output context.

The benchmark processes text in overlapping 510-token windows with a stride of 384 tokens, then uses max pooling to address truncation. For the leaderboard, classifiers use a 0.5 threshold on the injection class, except LLM Guard, which uses its shipped defaults. Catches are attacks correctly flagged; false positives are benign outputs wrongly flagged. Latency is median per-call CPU time.

The count of 629 attacks also appears in the AgentDojo paper, but in a different evaluation: the paper describes 629 security test cases across 97 user tasks and evaluates agents, attacks, and defenses in its own environment. Those agent-level outcomes are not the same as the text-classification scores in this detector benchmark. See the AgentDojo paper at NeurIPS 2024 for that context.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Default results: catch rate must be weighed against false alarms

The results below are from the repository’s default configuration. The three numeric columns represent embedded attacks caught, benign outputs flagged, and distinct attacks caught when tested alone.

Detector Embedded attacks caught Benign outputs flagged Attacks caught alone Median CPU latency
jailbreak-detector-large 319/629 (51%) 2/97 (2%) 25/27 110 ms
protectai-deberta-v2 145/629 (23%) 4/97 (4%) 27/27 163 ms
llm-guard (shipped threshold 0.92) 124/629 (20%) 2/97 (2%) 27/27 124 ms
prompt-guard-2-86m 6/629 (1%) 0/97 (0%) 0/27 149 ms
prompt-guard-2-22m 0/629 (0%) 0/97 (0%) 0/27 55 ms
Regex baseline 0/629 (0%) 0/97 (0%) 0/27 0.05 ms
preamble-defense 556/629 (88%) 46/97 (47%) 26/27 124 ms
testsavant-defender 370/629 (59%) 47/97 (48%) 15/27 37 ms
deepset-deberta 629/629 (100%) 95/97 (98%) 27/27 146 ms
fmops-distilbert 629/629 (100%) 95/97 (98%) 27/27 31 ms

DeBERTa by deepset and fmops-DistilBERT caught every embedded attack in this set, but each also flagged 95 of the 97 benign outputs. A detector that blocks that much normal content may be unusable in a real workflow. By contrast, jailbreak-detector-large caught fewer attacks, but its false-positive count was much lower. The useful comparison is therefore not simply “which detector caught the most?” but “how many attacks did it catch at an acceptable false-alarm rate?”

Why testing attacks in context matters

Two detectors recognized all 27 distinct attack texts when those strings were scored alone, yet caught far fewer when the same attacks appeared inside tool output: ProtectAI DeBERTa v2 caught 23% of embedded attacks, and LLM Guard caught 20%. This gap shows why a score on isolated injection strings cannot stand in for performance on mixed content. A detector may respond differently when malicious instructions are surrounded by legitimate-looking text.

Input handling also matters. The benchmark’s windowing and max-pooling approach is one specific way to score long outputs; these results do not establish how every deployment’s truncation, chunking, or aggregation choices would perform.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Threshold calibration changes the result—and its meaning

Prompt Guard 2 86M caught just 6 of 629 attacks (1%) at the default benchmark setting. In a separate fixed-false-alarm-budget experiment, the repository reports 621 of 629 pooled catches (99%) at a threshold of 0.003, with 5 of 97 unseen benign examples flagged (5%). The calibration used three AgentDojo domains and evaluated the remaining domain; the reported fold catches were 97%, 100%, 100%, and 100%, with a minimum-fold estimate of 97% (95% confidence interval 94–98%).

That improvement is not evidence that Prompt Guard 2 “solves” prompt injection. The repository warns that all AgentDojo attacks in this set share one wrapper template, so the tuned detector may be recognizing that template rather than generalizing to other attacker wording. The held-out-domain evaluation is within this benchmark, not an external validation across different attack styles or production traffic. The 97 benign examples also make the false-alarm estimate sensitive to individual cases.

Calibration did not produce uniformly strong held-out results for every model. In the same repository analysis, fmops had 48% pooled catches and a 26% minimum fold; jailbreak-detector-large had 51% pooled catches and a 17% minimum fold; and deepset had 0% pooled catches and a 0% minimum fold. The source cautions that intervals overlap and close rankings should not be over-interpreted at this sample size.

What these numbers can—and cannot—tell you

The benchmark measures how specific detector implementations classify the benchmark’s embedded attack and benign text, under the stated default thresholds and a separate cross-domain calibration procedure. It does not run a live agent. It therefore does not show whether a model would follow or ignore a detected instruction, whether an agent would make an unsafe tool call, or whether an authorization policy would block that action.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That distinction matters operationally: suspicious text can be flagged while an unauthorized tool call is plainly worded, and a detector can block ordinary work through false positives. A text score is not an authorization decision. Shastri’s proposed engineering direction is to calibrate against the traffic a deployment actually sees and enforce policy using the action and the provenance of its arguments, rather than relying on a text score alone. That is the benchmark author’s recommendation, not a result established by the leaderboard.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to read a detector comparison

For a practical evaluation, compare more than the headline catch rate:

  • Embedded attack catch rate: Does the detector find attacks inside the kind of tool output it will process?
  • False-positive rate: How often does it block benign content, and would that disrupt normal work?
  • Held-out performance: Does a tuned threshold carry over to a domain or traffic slice not used for calibration?
  • Context and windowing: How are long inputs split and scored, and does surrounding text change detection?
  • Latency: What is the measured per-call time under the benchmark’s CPU setup, and does the deployment’s hardware and workload make that number relevant?

The repository reports default catches, false positives, context-free catches, and median CPU latency, plus a held-out-domain calibration experiment. It does not report live-agent outcomes or a production operational-cost study, so neither should be inferred from these figures.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.