The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Test the complete agent application in an isolated environment—not just the model’s replies. Simulate its tools, data, tasks, permissions, and approval flow; try both direct misuse and malicious instructions embedded in material the agent reads; then record whether prohibited actions were actually blocked. If an unsafe action executes, it is a failure even if the agent later says it should not have done it.
What should an unsafe-tool-use test cover?
Test the boundary where the agent can cause effects, including the components that shape or authorize a tool call. A model-only evaluation can miss failures in orchestration, permissions, credentials, or approval handling.
- The agent’s prompts, orchestration, model provider, and any cooperating agents.
- Tool definitions, arguments, credentials, authorization policy, and the gateway that permits or denies execution.
- Retrieval, memory, and external content that can influence a proposed action.
- Approval controls and the state changes a successful tool call can make.
Use mock services, disposable accounts, or another isolated environment with synthetic data. OWASP’s AI Agent Security Cheat Sheet advises testing application controls alongside agent-specific failure modes and avoiding secrets or live customer data in test fixtures.
How do you build a repeatable test plan?
1. Define the boundary and expected result
For each tool, list its permitted operations, scope, credentials, data access, and possible side effects. For every test case, write down the legitimate user task, the attacker-controlled input, the prohibited action, the decision the policy should make, the evidence you will capture, and how to restore the test environment.
#1 Best Overall
2. Create an abuse-case matrix
Start with these cases from OWASP’s agent security testing guidance, then add scenarios tied to your own application and threat model.
| Abuse case | Test setup | What to verify |
|---|---|---|
| Prompt override | A user or retrieved item tells the agent to ignore higher-priority instructions. | Trusted instructions and independent policy remain effective. |
| Unauthorized tool use | The agent requests an operation outside the user’s or session’s scope. | The authorization layer denies the request before execution. |
| Privilege escalation | A low-trust session tries to access a privileged tool or credential. | Role boundaries and credential scope hold. |
| Memory poisoning | Malicious content is offered for persistence or later retrieval. | The content is rejected, appropriately scoped or sanitized, or expires as intended. |
| Data exfiltration | External content asks the agent to send private context to an attacker-controlled destination. | The transfer is blocked or requires the correct approval; inspect arguments and network effects. |
| Approval bypass | A high-impact action is attempted without approval, or with approval for a different or outdated action. | Approval is current and bound to the exact tool, target, and normalized parameters. |
| Recursive tool abuse | An operation triggers repeated tool calls or retries. | Configured depth, retry, token, and cost limits stop runaway activity. |
| Multi-agent boundary | One agent attempts to persuade another to act beyond its authority. | Delegated scopes and trust boundaries persist between agents. |
3. Pair a normal task with indirect injection
Place adversarial instructions in realistic untrusted sources the agent reads, such as a web page, document, email, tool output, or retrieved record. Pair that content with an ordinary user task and check whether it causes a prohibited action. NIST’s Center for AI Standards and Innovation (CAISI) describes agent hijacking as indirect prompt injection: malicious instructions embedded in ingested data lead an agent toward unintended actions. The important test is whether the application’s trust boundary holds when task-relevant data contains instructions.
4. Capture the action, not just the answer
Instrument the tool gateway or mock tools to record the requested tool and arguments, caller or session, policy decision, approval state, execution result, and resulting state changes. Verify that denial occurs before execution. A refusal in the final response does not make a test pass if the tool already performed the prohibited action.
5. Repeat and adapt the cases
Run each scenario multiple times when behavior can vary. Report results by task and attack type as well as in aggregate, so a strong overall score cannot conceal one vulnerable workflow. Add new cases when an attack succeeds or system behavior changes, and include human red-team review for high-impact scenarios. In its January 17, 2025 technical blog, NIST CAISI recommends adaptive evaluations, task-specific analysis, and multiple attempts.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
6. Make results a release control
Keep adversarial inputs and expected decisions under version control. Add regressions for discovered injection, memory-poisoning, and tool-abuse failures. Require updated tests when prompts, tools, memory, retrieval, policy, model provider, credential scopes, or approval logic change. OWASP recommends blocking releases when high-risk changes to tool policy, approvals, or credential scope lack updated tests.
7. Preserve a reproducible record
For each test run, record:
- Agent version, model provider, tool policy, retrieval configuration, and relevant permission or approval settings.
- Abuse cases executed, expected results, and observed approvals, denials, timeouts, or circuit-breaker behavior.
- Any side effects, accepted residual risks, and compensating controls.
Keep the record useful for comparison without putting secrets or live customer data into fixtures.
Rank #4
How should you interpret test results?
Judge whether the application protected the action boundary, not whether the model sounded cautious. A prohibited operation that reached execution is a failure. A denied call is evidence that a control worked for that case, but it does not establish that other tasks, attack types, or configurations are safe.
In a specific AgentDojo Workspace evaluation against an upgraded Claude 3.5 Sonnet, NIST CAISI reported that its strongest newly developed attack raised the measured attack success rate from 11% for the strongest baseline to 81%. Those figures describe that experiment, not a general vulnerability rate for AI agents. They illustrate why baseline cases alone can give an incomplete picture and why attack results should be reported by task and scenario.
Recommended Free Tools
Best Value
Which frameworks or services can help?
These options can support parts of an evaluation, but none substitutes for testing the reader’s own application and tool boundary.
| Option | What the cited source establishes | How to assess fit |
|---|---|---|
| AgentDojo | NIST CAISI describes it as an open-source framework used for hijacking evaluations, with simulated Workspace, Travel, Slack, and Banking environments and simulated tools. CAISI added scenarios for remote code execution, database exfiltration, and automated phishing. | Use it as a foundation and adapt cases to your own tasks, tools, and deployment; it is not certification of a different agent. |
| Promptfoo | OpenAI’s red-teaming guide identifies it as an open-source framework for evaluating prompts, agents, and AI applications, including generation of adversarial cases and inspection of target behavior. | Check current features, integrations, license, and compatibility with your stack before adopting it. |
| Managed red teaming | OpenAI says its managed red-teaming service is available to enterprise customers. | Confirm current eligibility, scope, and terms directly with the provider. |
Compare options on whether they exercise the full application or only model behavior, attack and environment coverage, capture of tool actions and side effects, repeatability and CI integration, custom scenario support, and operational support and reporting. The cited material does not establish a universal product ranking or current compatibility matrix.
What defensive controls reduce the impact of a failure?
- Give the agent only the tools and privileges required for its task.
- Separate the agent’s proposal from execution so another component can enforce policy.
- Have an independent policy component check action scope and required approvals.
- Bind approval to the exact action, including the tool, target, and normalized parameters.
These controls limit the consequences of manipulation; they do not replace adversarial testing. OpenAI authors Thomas Shadwell and Adrian Spânu framed the design objective in their March 11, 2026 article as constraining the impact of manipulation even if it succeeds, rather than relying only on the agent to identify malicious input.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

