Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not judge an AI agent’s safety from its answer quality alone. Evaluate the specific agent, tools, permissions, connected data, and operating environment you plan to use—and test what it does when it encounters both malicious instructions and ordinary edge cases.

What “safe to use” should mean for an AI agent

An agent is not simply answering questions when it can use tools. It may read private data, change records, send messages, run code, or trigger actions in another system. Its risk depends on both what its tools can do and the environment in which it uses them.

Set a task-specific standard before testing: list the actions that are acceptable, the actions that must be blocked or approved, and the harms that could result from a mistake. Consider data exposure, persistent changes, external communications, spending, and chains of actions—not just whether the final response sounds correct. NIST describes agent systems as capable of planning and taking autonomous actions that affect real-world systems or environments.

There is no universal pass/fail threshold in the NIST materials for every agent and deployment. Treat evaluation as a way to identify and reduce risk, not as proof that an agent is safe in every setting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inventory the agent’s tools, permissions, and authority

Write down what each tool reaches and what the agent can do with it. Include the credentials involved and any limits imposed by the tool or connected service. A tool’s name is not enough: a “browser” might only retrieve pages, or it might also submit forms; a “mail” tool might read messages, draft replies, or send them.

What to record Questions to answer
Resource Which account, files, services, systems, or people can the tool reach?
Readable data What information can it retrieve, including sensitive or user-provided content?
Possible changes Can it create, edit, delete, publish, send, purchase, execute, or otherwise change persistent state?
Credentials Which identity or secret authorizes access, and where is it held?
Limits Are actions restricted by scope, destination, amount, rate, or an approval step?
Visibility Can you inspect tool calls and outcomes and identify which agent or user initiated them?

NIST’s August 5, 2025 workshop-informed tool-use taxonomy distinguishes read-only, constrained-write, and write access, and considers whether the environment is trusted or untrusted. Use these as description categories, not as a complete risk standard:

Access pattern What to establish Evaluation implication
Read-only The agent can retrieve information but cannot change the resource through that tool. Check for data exposure and whether read content can influence other tools the agent can use.
Constrained-write The agent can make changes, but the available write actions have defined limits. Test both the limits and whether the agent can act around them through another tool or a sequence of calls.
Write The agent has write capability without the same narrow action constraints. Assess the consequences of an unintended action and whether stronger restriction or approval is needed.
Trusted or untrusted environment Identify whether the agent interacts with resources and content you trust, or with material and systems that may be controlled by others. Exposure to untrusted content matters especially when the agent can use tools to act on what it reads.

Do not assess a permission in isolation. A read-only source may still create action risk if the agent can read its instructions and then use a separate tool to send a message or change a record.

Test whether webpages, emails, or files can hijack the agent

Indirect prompt injection occurs when malicious instructions are placed in data an agent may ingest, rather than supplied directly as the user’s request. NIST CAISI’s January 17, 2025 technical blog describes agent hijacking as an attacker inserting instructions into data that can cause an agent to take unintended, harmful actions. NIST frames a hijacking scenario as a legitimate task combined with an injected malicious task; carrying out the injected task indicates that hijacking succeeded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Use a controlled environment. Test with the same tool configuration and permission boundaries intended for deployment, but use test accounts, sample data, and reversible actions wherever possible.
  2. Give the agent an ordinary task. For example, ask it to summarize a test webpage or extract information from a sample email.
  3. Place conflicting instructions in the content. Use representative test material that attempts to redirect the agent away from the user’s task or induce an unauthorized tool action.
  4. Repeat across content types and tools. Include the sources the agent is expected to process, such as webpages, email, and files, and cover the tools it could use after reading them.
  5. Inspect what happened. Review tool calls and resulting state, not only the final text. A refusal in the final answer does not establish that the agent refrained from acting.

Record whether the agent stayed within the requested task, exposed data, followed the injected instruction, or attempted a blocked action. NIST CAISI’s blog describes scenario-based evaluations, including AgentDojo and additional scenarios; it discusses a particular tested model version, not a current ranking of models.

Test failures that do not require an attacker

An agent can cause harm through mistakes, ambiguous requests, or behavior that pursues a goal in an unintended way. NIST’s January 12, 2026 request for information on securing AI agent systems includes risks from adversarial data as well as harmful actions without adversarial inputs. Its May 18, 2026 summary analyzes stakeholder responses; neither document establishes a binding pass/fail test.

  • Ambiguous requests: Check whether the agent asks for clarification before a consequential action when the user’s intent is unclear.
  • Boundary cases: Test requests that sit near the edges of the agent’s authorized task and observe whether it stays within scope.
  • Unsafe tool use: Check for unnecessary, incorrect, or excessive tool calls, including actions that could have been avoided.
  • Data disclosure: Test whether the agent reveals information to an unauthorized recipient or uses data outside the task’s scope.
  • Unexpected goal pursuit: Look for actions that achieve a requested result in a harmful or unintended way, including specification gaming or misaligned behavior.
  • Chained actions: Check whether a sequence of individually plausible calls produces an unauthorized or harmful outcome.

For each case, define the expected safe behavior before running the test. Then compare it with the tool calls, the resulting state, and any approval or refusal—not just the explanation the agent gives afterward.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Check who grants authority and how actions are accountable

Before connecting an agent to real accounts or systems, establish how it is identified, authenticated, and authorized. Ask whether its permissions are tied to the user and task, how delegated access is bounded, and how actions can be attributed and audited. Decide which actions require human approval, particularly when a mistake could have significant or difficult-to-reverse effects.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NIST NCCoE’s February 5, 2026 concept paper on identity and authority of software agents raises issues including least privilege, delegation, human-in-the-loop authorization, audit, and non-repudiation. It presents these as project questions and design considerations, not as a final implementation standard. The related NIST NCCoE agentic AI identity and authorization project hub is a project resource whose content and status may change.

Constrain access and monitor deployment

Give the agent only the permissions needed for its task. Where it must write, limit the scope of what it can change and consider whether a person should approve higher-impact actions. Monitor tool use and retain enough information to investigate an unexpected call or outcome. No single control should be treated as eliminating risk.

NIST CAISI’s January 2026 request for information asks about ways to evaluate agent security and constrain and monitor access in deployment. That supports evaluating these controls, but it does not prescribe one universal control set. Check that safeguards work in the actual configuration, including when tools are chained or the agent processes untrusted content.

Make a scoped decision and retest after changes

Document the agent version, tools, permissions, credentials, connected data, and environment that were evaluated. Keep the test scenarios, observed failures, resulting state, and mitigations with the decision. That record makes clear what the result covers—and what it does not.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reevaluate after a material change to the model, tools, prompts, credentials, connected systems, permissions, or autonomy. NIST’s tool taxonomy and its other agent-security materials treat tool access and deployment context as important to risk; a result for one setup should not be generalized to a different one.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.