Approval-required outcomes are not hard blocks. In a recorded RedCode evaluation, 713 of 720 in-scope attack cases were either blocked or sent for approval—but 124 of those 713 still required a human operator to decide. A useful security benchmark makes that distinction, its test scope, benign-user friction, method, and date visible beside the headline result.
What the approval split changes in a security result
Security tools can produce materially different outcomes: block an action, allow it only after an operator approves it, or pass it without intervention. Combining blocks and approval requests into one “stopped” figure may describe actions that did not proceed automatically, but it hides who made the final security decision.
In the September 4 recorded RedCode run described by Alan Fu, deterministic rules evaluated 1,410 attack records. The run excluded 690 as outside its declared threat model, leaving 720 in-scope cases. Of those, 589 returned BLOCK, 124 returned AUTH, and seven returned PASS. Thus 713 cases were blocked or required approval; only 589 were hard-blocked. The 124 AUTH results depended on an operator response, so they should not be represented as blocks. Fu’s account of the RedCode run provides the reported counts and their context.
The denominator matters just as much as the outcome labels. Reporting “589 of 1,410 blocked” would count cases the evaluation deemed outside scope; reporting only “713 stopped” would obscure the approval dependency. Give readers the in-scope denominator and the excluded count, then show the outcome categories separately.
#1 Best Overall
Put benign friction beside attack outcomes
A guardrail can intervene on benign activity as well as malicious activity. In the same run, 60 synthetic benign controls produced 56 PASS outcomes, three AUTH outcomes, and one BLOCK. Those four interventions are useful friction data, but the controls were synthetic—not production user sessions. The run report identifies these counts and the nature of the controls.
For a fuller picture, publish benign outcomes alongside attack outcomes and explain how benign examples were selected and labeled. A benchmark that reports attack catches without showing benign interruptions leaves readers unable to assess the usability cost or interpret a false-positive claim.
Rank #2
Read narrow results as narrow results
Subcategory results can clarify what a test actually exercised, but they do not establish universal protection. In the recorded run, all 30 reverse-shell-listener cases returned BLOCK. That supports a claim about those cases in that evaluation; it does not demonstrate detection of every reverse shell.
For 60 process-kill cases, all required intervention: 13 returned BLOCK and 47 returned AUTH. That split shows why “intervention” and “block” are not interchangeable labels. Preserve the categories when presenting any subcategory totals.
Free tools Windows power users keep installed
One-click scans. No signup required.
Check what the evaluation did—and did not—test
The RedCode evaluation replayed mapped tool-call cases through a deterministic engine. It did not run a live model through a complete attack campaign and did not measure the full adaptive layer. Its counts are historical recorded results, not a fresh test of whatever release a reader encounters later. The run description and the author’s discussion of test-linked guarantees explain those limits.
Before comparing two benchmark claims, check whether they evaluate the same kind of evidence. A replay may be reproducible but not capture live adaptive behavior; a detection score is not necessarily a governance audit; and a test tied to one host or release does not automatically describe another. Fu also points to a host-parity matrix in the article, underscoring that host-specific evidence should be matched to the host being evaluated.
Rank #4
Evaluate benchmark quality, not just the score
Look at corpus construction and labels
OASB describes 222 standardized attack scenarios with mappings to MITRE ATLAS and OWASP. Its documentation describes adapters running against a suite and marking undeclared capabilities N/A rather than FAIL. Its specifications distinguish a tool-detection benchmark from governance auditing. These are useful structural details, but they are not evidence that a particular product has passed. See the OASB project page, OASB version 0.4.0 specifications, and OASB-1 getting-started documentation.
Label provenance can change what a metric means. OASB disclosed that it withdrew F1, precision, and false-positive-rate figures after finding that its benign class had been selected using the scanner’s own labels, making the near-zero false-positive result circular. The project reports recall of 223/270 (82.6%) on author-created attack fixtures, and 234/495 (47.3%) when self-labeled samples are included; it says it is remeasuring with corpora it neither owns nor labeled. Those figures describe particular corpus compositions, not a general rate for agent-security tools. The OASB page describes the disclosure and remeasurement.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Ask who ran the test and whether it stayed held out
MoorAI reports three scored runs, all executed by its maintainer, and says its repository has no third-party lab reproductions. Its methodology describes locked test halves intended to preserve a check against tuning. That makes the distinction between maintainer-run results and independent validation important; a “held-out” split is meaningful only if it remained unseen during development. Details are on MoorAI’s benchmark methodology and results page.
Treat proposed frameworks as proposals
A July 5, 2026 IETF Internet-Draft proposes four first-level dimensions and 55 second-level metrics across static, dynamic, attack-defense, compliance, and quantitative evaluation. It is an informational draft, not a certification or a product result. If cited, identify it as a proposal and retain its date and status: IETF Internet-Draft: Security Evaluation Benchmark for AI Agents.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A checklist for reporting or comparing a benchmark
- Product and scope: Name the product, version, host, corpus, threat model, and evaluation date. State how many cases were included and excluded, and why.
- Outcome definitions: Publish counts for BLOCK, AUTH or approval-required, PASS, and any other result. Do not fold approval into hard blocks.
- Benign friction: Report benign controls and interventions beside attack outcomes. Identify whether controls are synthetic or production-derived and explain label provenance.
- Evaluation method: Say whether cases were replayed or run live, whether a model was involved, and whether adaptive behavior was evaluated.
- Independence and generalization: Disclose who ran the evaluation, whether an independent party reproduced it, and whether a held-out set remained untouched during development.
- Claim-to-test fit: Match the evidence to the property being claimed and the release and host where it is expected to apply.
These distinctions are not bookkeeping. They determine whether a headline describes a block, a human checkpoint, a narrow corpus result, or evidence for the product and environment a reader actually cares about.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

