Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To investigate an AI agent’s unauthorized tool call, preserve the available records, reconstruct the full path from request to downstream effect, and verify the action in the system that owns the affected resource. An agent transcript can show what was proposed or recorded; by itself, it may not establish what actually executed or changed.

1. Preserve evidence and define the incident window

Start by recording when the behavior was detected, the suspected run or session, the identities involved, the systems potentially affected, and whether activity is still ongoing. Preserve the records most likely to be overwritten or removed by routine retention or cleanup.

Collect what is available from the agent, tool executor, identity and authorization systems, downstream services, and the configuration or policy versions in effect at the time. OWASP recommends audit trails for agent decisions and actions, including detailed records of tool invocations, context changes, and user-agent interactions in its AI Agent Security Cheat Sheet and MCP Top 10.

Record gaps as gaps. Not every agent deployment exposes the same events, and the absence of a trace is not proof that an action did not happen.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Reconstruct the action chain

Build a timeline that connects the initiating human or service principal to the agent session, any recorded model turn, the proposed tool call, the executor’s decision, the execution result, and the downstream event or changed state. Preserve original timestamps and identifiers so that records from different systems can be correlated.

For each event, capture these fields when available:

  • Timestamp, run or session ID, agent identity, and initiating principal.
  • Tool name, normalized arguments, and target resource.
  • Authorization result, applicable approval ID, and executor response.
  • Downstream event ID, result status, or observed resource state.

OWASP recommends structured decision metadata and tool-call outcomes; its MCP guidance warns that limited telemetry can impede investigation. If a field or event is missing, note which source was checked and what it could not establish rather than filling the gap with an assumption.

3. Verify what actually executed

Check the tool server or executor and the downstream service that owns the affected resource. A transcript or model-generated request can document what the agent proposed or what its layer recorded. An accepted request, a successful tool response, and a verified lasting change are different findings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, a tool response may report success while the downstream state remains unchanged, or a downstream audit record may show an operation that is absent from the agent trace. Compare the relevant records and identify what each one proves. NIST notes that tool observability varies: some actions can be seen in existing logs or transcripts, while others require additional ways to observe their effects. See NIST’s lessons from the consortium on tool use in agent systems.

4. Determine whether the call exceeded authority

Judge the action against the original task and the permissions and controls that were effective when it occurred. A tool call may be technically permitted by a credential yet still exceed the user’s intended task or an approval requirement.

For the specific call, establish:

  • Which tool functions were enabled and which resource or resources they could reach.
  • Whether authority was read-only, constrained-write, or broader write access.
  • Which human, service, or delegated identity the downstream system saw.
  • What credential scope and policy version applied at the time.
  • Whether approval was required for this action and, if so, whether it was granted for the specific operation.
  • Whether the downstream system independently enforced authorization or relied on the agent to decide.

OWASP’s LLM06:2025 Excessive Agency identifies excessive functionality, permissions, and autonomy as common sources of risk. NIST also recommends evaluating access patterns and tool constraints. A call can be outside authority even if it was successfully executed; determine both the control decision and the actual effect.

5. Test plausible causes against the evidence

Do not assume that the model alone caused the incident. Check hypotheses against the records and configuration you preserved, including:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Unnecessary, open-ended, or overly powerful tool functions.
  • Overbroad, stale, or incorrectly delegated credentials.
  • Missing or ineffective authorization checks in the downstream service.
  • An approval control that was absent, misconfigured, or bypassed.
  • Direct or indirect prompt injection that influenced tool use.
  • Misleading or compromised output from an extension or tool.
  • Multi-agent delegation that changed which identity acted or what scope it could exercise.

OWASP describes both excessive tool functionality, permissions, or autonomy and unexpected or manipulated inputs as paths to harmful action. These are investigation hypotheses, not conclusions about any particular incident; corroborate them with artifacts such as policy versions, executor decisions, identity records, and relevant context or tool output.

6. Assess impact, persistence, and reversibility

For every verified action, identify the resources accessed, data read or exposed, state changed, and any external communication or transaction. Look for persistence and possible follow-on actions as well as the immediate effect.

Classify the capability and observed effect separately. A read-only permission may still expose sensitive information; a write-capable tool may or may not have changed state in this incident.

Authority or effect What to establish
Read-only Which resources or data were accessed, whether the data was sensitive, and whether it was sent or exposed elsewhere.
Constrained write Which limited changes were possible, whether a change occurred, and whether the affected state can be safely restored.
Unrestricted write Which resources could be changed, what actually changed, whether effects persist, and what further actions may have followed.

Assess severity in the context of the environment, the tool’s permissions, the statefulness of the action, and whether reversal is safe and authorized. NIST’s tool-use guidance treats these as complementary dimensions: capability alone does not establish impact, and impact cannot be inferred from the agent’s response alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

7. Contain the unsafe path

Contain the specific route that enabled further unauthorized activity. Depending on the system and verified impact, that may mean pausing or restricting the affected agent or tool, revoking or narrowing credentials, or blocking unsafe downstream operations. Preserve the records needed to understand the incident while doing so.

Choose containment measures for the affected service and risk. Shutting down an entire system indiscriminately can disrupt unrelated work, while leaving a live unsafe capability available can allow additional actions. OWASP recommends minimizing extensions and permissions, enforcing authorization downstream, and requiring human approval for high-impact actions.

8. Recover and close the control gap

Reconcile the affected downstream state against authoritative records. Reverse changes only when the correction is safe and authorized, and restore a capability only after the control that failed has been addressed.

Use the findings to reduce the chance of recurrence: remove unneeded tool functions, scope permissions to the task and user, enforce authorization independently for each downstream operation, and require action-specific approval for high-impact operations. Improve monitoring and alerting for the events needed to reconstruct future calls. OWASP’s AI Agent Security Cheat Sheet covers least privilege, approval, audit trails, and monitoring; NIST emphasizes matching observability to the tools and environment actually deployed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What incomplete telemetry means

OWASP MCP Top 10, MCP8:2025, states: “Limited telemetry from MCP servers and agents impedes investigation and incident response.” If logs do not capture a needed link in the action chain, state that limitation explicitly and use executor or downstream records where available. The public evidence may establish that a call was proposed without establishing whether it ran or what it changed; keep those conclusions separate.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.