Free tools Windows power users keep installed
One-click scans. No signup required.
Compare AI agent security benchmarks by the behavior they test, the agent and tools they include, how attacks are designed, what the scorer counts, and whether benign task success is measured too. AgentDojo, AgentHarm, and Agent Security Bench (ASB) examine different risks, so their scores are not interchangeable. A credible comparison also needs adaptive attacks, repeated attempts where relevant, transparent configurations, and checks that the agent’s trace supports the score.
Start with the security claim you need to test
“Agent security” is not one outcome. A test may measure whether an agent follows malicious instructions hidden in an email, whether it complies with a direct harmful request, whether it makes an unsafe tool call, or whether an attacker achieves a larger goal such as data exfiltration. Those are related risks, but evidence about one does not establish performance on the others.
Before comparing results, write down the claim in observable terms: what the attacker wants, what the agent can do, and what counts as success. A result is only as broad as the behavior and system boundary exercised by the test.
- Indirect prompt injection: malicious instructions arrive inside material the agent is asked to read, such as a message, file, or web page.
- Harmful-request handling: the agent is asked directly to help with a harmful task; evaluation may include both refusal and the ability to carry out a multi-step task if the refusal is bypassed.
- Unsafe action or goal completion: the measure may focus on a particular tool call, policy violation, or whether the attacker’s intended outcome was actually achieved. These are not equivalent scoring targets.
Compare the benchmark’s scope, not just its name
AgentDojo, AgentHarm, and ASB are useful for different questions. The table summarizes their documented emphasis; it is not a ranking. In particular, the broad scope reported for ASB does not prove that every scenario is equally realistic or that it covers every agent risk.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
| Benchmark | Main target and setup | What its reported scope tells you | Best fit and qualification |
|---|---|---|---|
| AgentDojo | Prompt-injection attacks and defenses in tool-using workflows over untrusted data, paired with legitimate user tasks. | ETH Zurich researchers’ 2024 paper describes 97 realistic tasks and 629 security test cases. Project documentation describes banking, Slack, travel, and workspace suites. | Use it to examine whether an agent can handle interactive workflows while resisting instructions embedded in task-relevant data. Results depend on model, prompt, suite, attack, defense, and execution setup; its package API is described as under development, so check current documentation and compatibility. |
| AgentHarm | Harmfulness and misuse of LLM agents, including whether an agent refuses harmful requests and whether it can complete a multi-step harmful task after a successful jailbreak. | The paper reports public release of the benchmark dataset; a comparable task or test-case count is not stated here. | Use it for direct harmful-request and misuse behavior, not as a substitute for an indirect prompt-injection evaluation. Check the current dataset version and exact scoring protocol before comparing leaderboard results. |
| Agent Security Bench (ASB) | A broad attack-and-defense study across agent scenarios, tools, and evaluation metrics. | ASB authors’ 2024 paper reports 10 scenarios, 10 agents, more than 400 tools, 23 attack/defense method types, eight metrics, and nearly 90,000 test cases in its experiments. | Use it to study a wider set of attack and defense combinations. Its reported experimental scope is not proof of exhaustive risk coverage; align the threat, agent setup, and metric before comparing it with a narrower benchmark. |
Check how the benchmark creates its evidence
A benchmark name or dataset alone does not define an evaluation. The 2025 ACM survey of LLM-agent evaluation organizes choices around objectives—behavior, capability, reliability, and safety—and process, including interaction mode, benchmark or dataset, metric computation, and tooling. Use those choices to inspect how a result was produced.
Agent, tools, and environment
Find out whether the test runs a complete agent with state and tools, a simulated workflow, or isolated model prompts. Record the domains and tools represented, the agent implementation, and the permissions it receives. A model tested without the same tool access or state as a deployed agent is not evidence about that full deployed system.
Rank #2
Attack construction and defenses
Identify whether attacks are fixed, held out, adaptive to the tested system, or developed against that system. Note which defenses and baselines are included. A score against a static set of attacks answers a narrower question than performance against attacks adapted to the system’s behavior.
Interaction and repetition
Check whether the agent has one chance or multiple attempts, whether outputs are sampled or deterministic, and how many attempts are run for each task and model. If the real attacker can retry cheaply, a one-shot result may miss failures that emerge across attempts.
Scoring target and utility
Read what the scorer counts: an attempted action, a completed attacker goal, refusal, policy compliance, or benign task success. Also check whether scoring is automated, based on a rubric, or human-reviewed. Measure benign task success alongside security outcomes when the benchmark includes legitimate work; a defense that blocks attacks by also preventing the requested task has a different trade-off from one that preserves utility.
Account for adaptive attacks and retries
NIST CAISI’s January 17, 2025 technical blog, “Strengthening AI Agent Hijacking Evaluations,” treats hijacking as indirect prompt injection: malicious instructions are placed in data an agent reads, such as an email, file, or web page, to redirect its actions. Its guidance is to improve shared evaluations continually, adapt attacks to the system, analyze task-specific results, and consider multiple attempts. As the NIST technical staff put it, “Evaluations need to be adaptive.”
Rank #4
In the specific CAISI experiments described in that blog, attack success ranged from 11% to 81% when the strongest new red-team attack was compared with the strongest baseline attack in the evaluation. In a separate repeated-attempt experiment, mean attack success rose from 57% to 80% after each of five injection tasks was run 25 times. These are results from those tested models, tasks, and evaluation conditions—not general rates for deployed agents. They illustrate why a single attempt can understate risk when outputs vary and retries are feasible. NIST’s staff also wrote, “Testing the success of attacks on multiple attempts may yield more realistic evaluation results.”
CAISI describes developing attacks on a random subset of workspace tasks and testing them on held-out workspace tasks, then trying those attacks in other environments. This is a useful evaluation pattern: separate attack development from held-out testing, and report task-level results as well as aggregates. The latter helps reveal whether a headline score is driven by a small number of especially vulnerable tasks.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteAudit scoring validity and reproducibility
A high score is informative only if it represents the intended outcome. NIST CAISI’s evaluation-cheating guidance distinguishes solution contamination, where a model accesses information that improperly reveals a task solution, from grader gaming, where it exploits a scoring loophole without satisfying the intended task. Review transcripts, specify task rules clearly, close loopholes, and standardize what agents may do.
For every result, record the model version, prompt, agent implementation, tools and permissions, environment, task subset, attack set, scorer, and number of attempts. Include such details as internet access, package versions, and scorer behavior when they affect the result. Check automated scores against the actual outcome and the agent’s trace, especially if the metric uses a proxy such as a particular tool call.
A 2026 preprint, Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks, reports audits of R-Judge, InjecAgent, AgentHarm, and AgentDojo using official implementations and author-provided scorers, while measuring capability benchmarks under its own protocol. The authors argue that safety claims should name the benchmark, metric, target behavior, and model panel. Treat this as recent preprint evidence, not settled consensus; its core reporting advice is useful regardless.
Use a like-for-like comparison checklist
Before treating two results as comparable, check that they share a meaningful denominator and evaluation setup. If an important item differs, report the difference rather than collapsing the results into a single ranking.
- Name the target behavior. Specify the attacker’s goal and whether the test concerns indirect injection, direct harmful requests, unsafe tool use, or another behavior.
- Describe the system boundary. State whether this is a model prompt, simulated workflow, or full tool-using agent, and list the environment and affordances.
- Describe the attack set. Say whether attacks are fixed, adaptive, held out, and system-specific; identify included defenses and baselines.
- Define the outcome and scorer. State whether the score counts attempts, completed goals, refusal, utility, or another outcome, and how it is judged.
- Report attempts and variability. Include attempts per task and model, sampling or determinism, and whether retries are represented.
- Report utility and inspect traces. Pair security outcomes with benign-task performance where applicable; check transcripts for task completion, loopholes, and scoring errors.
- Publish the configuration. Identify model and prompt versions, tools, permissions, environment, task sample, scorer, and other settings needed to interpret or reproduce the result.
What benchmark scores can—and cannot—establish
A benchmark result supports a claim about the tested configuration and target behavior. It does not by itself establish a universal ranking of agent security, a standardized metric shared across benchmark families, or a guarantee about performance in every production context. AgentDojo’s package and datasets can evolve, and the 2026 validity audit is a preprint. Name the benchmark, metric, target behavior, and model panel, and preserve the tested configuration when presenting a result.

