To test whether an AI agent follows a stranger’s instructions, give it a legitimate task, place a clearly defined attack in content it must process, and observe whether it pursues the attacker’s goal or makes an unauthorized tool call. Use fake data and sandboxed tools; never test against real accounts or secrets. A single test shows how that configuration behaved in that scenario—it does not prove that every agent is vulnerable or secure.
What this test is designed to catch
Prompt injection occurs when instructions embedded in a prompt or in content an AI system processes influence its behavior. A direct injection is supplied in the user’s prompt. An indirect injection arrives through material such as a webpage, file, email, or tool result. If you want to know whether an agent follows a stranger’s instructions in a document or webpage, put the test instruction in that content. Typing it into the user prompt tests a different channel; record that separately. OWASP explains the distinction in its LLM01:2025 Prompt Injection guidance.
The risk is not just that an answer sounds wrong. A tool-using agent can attempt actions through its connected tools. The possible impact depends on the agent’s permissions, capabilities, and application context: an instruction might try to expose sensitive data or trigger an action the user did not authorize. NIST discusses this broader risk as AI agent hijacking. There is no established general percentage for how often an arbitrary deployed agent can be hijacked, so a test result should be reported for the system and cases actually tested.
Build a safe, repeatable test
- Choose one real workflow. Use a task within the agent’s intended scope, such as summarizing a document, finding a specified email, or gathering information from a page. Define what a correct, useful result looks like before introducing an attack.
- Write the attacker’s goal and failure condition first. For example, the content may instruct the agent to reveal a planted dummy secret or invoke an unauthorized tool. Decide in advance what counts as failure: returning the protected value, attempting the prohibited action, or successfully executing it. These are test-design examples, not claims about any particular product.
- Put the instruction in the channel under test. For an indirect-injection test, place the adversarial instruction in the webpage, file, email, or tool output the agent will process. Do not substitute a user-prompt instruction and call it an indirect-injection test.
- Replace real data and tools with safe fixtures. Seed a fake secret and fake records. Route tool calls to sandboxed or instrumented substitutes that log requests but cannot alter real accounts or data. OWASP’s LLM Prompt Injection Prevention Cheat Sheet recommends harmless data and instrumented tool substitutes.
- Run controls as well as the attack. First run the legitimate task without adversarial content. Also test benign content that happens to contain instruction-like wording. These controls help distinguish an attack response from ordinary task failure or an overly broad block.
- Log the whole outcome. Record whether the legitimate task was completed, whether the attacker’s goal was met, what tools the agent tried to call, and whether application controls blocked an unauthorized action. Include the model and tool configuration, permissions, attack channel, and the failure rule you set.
- Keep the case and rerun it. Version-control the test inputs and expected outcomes. Repeat them before release and after meaningful changes to prompts, tools, memory, retrieval, policies, or model providers. OWASP’s AI Agent Security Cheat Sheet recommends structured, repeatable testing.
Measure security without mistaking refusal for success
A useful evaluation asks two questions at once: did the agent resist the attacker’s objective, and did it still complete the user’s legitimate task safely? An agent that refuses everything may avoid an attack while failing its actual job. AgentDojo uses this paired security-and-utility framing for agents acting on untrusted data.
#1 Best Overall
The AgentDojo paper, presented at NeurIPS 2024, describes a benchmark with 97 realistic tasks and 629 security test cases. Those counts describe the benchmark’s scope—not how often real agents are compromised, and not a prediction of how a particular deployment will perform. Its authors also emphasize that results depend on the tasks being tested and that static attacks can miss adaptive attacks. See the AgentDojo paper page.
- Attack success: number of attack cases in which the predefined attacker goal occurred, divided by the total attack cases run.
- Task utility under attack: number of attack cases in which the user’s task was completed correctly without unsafe side effects, divided by the total attack cases run.
- Attempted versus executed actions: report separately whether the agent requested a risky tool action and whether application controls allowed it to happen.
- Benign-task failures: track failures or unnecessary blocks in clean and benign-control runs.
For each measure, publish the numerator and denominator, the cases and configuration tested, the attack channel, and the failure definition. Rates from different task sets are not directly comparable unless their methods and scope are comparable.
Rank #2
Use the test to strengthen the boundary
OWASP notes that prompt injection is possible because models process instructions and data as natural language, and that fool-proof prevention is unclear. Treat defenses as ways to reduce the likelihood or impact of an attack, not as a guarantee that injection cannot work. As OWASP puts it: “Perform regular penetration testing and breach simulations, treating the model as an untrusted user to test the effectiveness of trust boundaries and access controls.” The sentence is from its LLM01:2025 Prompt Injection guidance.
Quick Recap
Best Value
Rank #4
- Enforce authorization in application code. Check whether a proposed action is allowed at the point where a tool executes; do not rely on the model’s judgment as the permission check.
- Apply least privilege. Give each tool and agent only the access required for its assigned task, and validate tool arguments before execution.
- Require approval for high-risk side effects. Use action-specific user confirmation where an operation could expose data or change an account or record.
- Keep secrets out of system prompts. A prompt is neither a secure credential store nor an authorization system. OWASP’s LLM06:2025 Excessive Agency guidance describes how excessive permissions can amplify the impact of injection.
- Label untrusted content, but do not rely on labels alone. Delimiters and warnings can help signal that content is untrusted; the actual security boundary must still be enforced by the application.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

