Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A passed run can show that a recorded execution path completed; it does not, by itself, prove the action was authorized. To make that case, connect the initiating identity and delegated authority to the exact action, a policy decision made before execution, any required approval, the runtime event, and evidence of the resulting effect.

What a successful run proves—and what it does not

A green status, agent transcript, or post-run log may document that a step completed. None alone establishes that a policy gate approved that specific action before it happened. The IETF Internet-Draft on agent-action evidence makes this distinction explicitly: a post-execution log documents what happened, but cannot prove that a policy gate existed beforehand. The draft is work in progress, not a finalized standard.

Authorization is a claim about authority and timing; execution is a claim about what ran. Keep the evidence for those claims separate, then connect them with identifiers and a clear chronology. Say exactly which parts your records establish. If you have only a transcript and runtime log, do not present them as proof of a prior authorization decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build an evidence chain for the specific action

A reviewer should be able to follow one action from its origin through its effect. Preserve the following evidence as distinct events, with timestamps and identifiers that let you correlate them.

1. Identify the initiator and delegated authority

Record the human or service that initiated the task, the agent identity, and the subject or role on whose behalf the agent acted. Preserve the granted scope and, where applicable, its validity period and revocation status. A run identifier alone does not establish who had authority to request or perform the action.

2. Bind authorization to the action that was proposed

Capture enough detail to distinguish the approved operation from a similar one: the operation, target tool, tool-schema version, resource, and arguments. Include the agent and represented subject identities. If your system uses an action digest, retain the digest and the action data needed to recompute it.

Microsoft’s Agent Governance Toolkit protocol describes this kind of binding and requires the execution boundary to recompute the digest from the action about to run. An approval for one binding cannot authorize another. If the tool, target, resource, schema, or arguments changed after approval, the earlier approval may not cover the executed action.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Preserve the policy decision before execution

Keep the decision identifier, policy rule and version, decision time, applicable scope, outcome, and reason where available. Record whether the policy returned allow, deny, or require_approval. A request waiting for approval is not the same as a denial, and neither outcome should be inferred from a later run status.

The key question is temporal: was the decision made before the operation crossed the execution boundary? A record created only after execution can describe the event, but cannot establish that the gate preceded it.

4. Attach any required approval to the same binding

When policy requires approval, preserve the approval request and its final disposition, the approver’s identity, decision time, and approved scope. Check that the approval covers the action that actually ran, rather than only the broader task or run. If the action changed, the approval record must still bind to the changed action for it to support authorization.

5. Correlate execution with an independently observable effect

Keep the runtime event, execution result, and any destination-system receipt or resulting artifact. Distinguish a model’s proposed action from an authorization decision, the command or tool call that executed, and the effect observed in the system it targeted. Northflank’s audit guidance uses these as separate categories; its platform logs or sandboxing do not, by themselves, establish that an agent’s business-level action was authorized.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use correlation IDs without making the agent the only witness

Assign a stable run or correlation ID and map it to native IDs from the orchestrator, policy engine, runtime, credential system, commands, artifacts, and destination platform. The mapping should let an auditor follow the action even when each component uses a different event ID.

Consider whether the agent or the same component being audited could omit, alter, or rewrite the records. Export evidence to access-controlled storage, monitor collection failures, and retain enough provenance to identify its source. A record is more useful when its origin and handling can be examined independently of the agent’s own account of the run.

Review a disputed run in this order

  1. Trace the request: identify the original task reference, initiator, agent, represented subject, and delegated scope.
  2. Match the action: compare the approved operation, tool, schema version, resource, and arguments or digest with the action recorded at the execution boundary.
  3. Verify the gate: inspect the policy version, decision outcome, timestamp, reason, and any required approval. Confirm that the decision preceded execution.
  4. Confirm the event and effect: correlate the runtime result with a receipt or artifact from the destination system where one is available.
  5. Assess the record set: determine whether the evidence is complete, independently retained, and clear about what it cannot establish.

This review is about whether the authorizing record covers the action that ran—not merely whether the records share a run ID. Redact secrets, restrict access, monitor for collection failures, and set retention by data class. Keep useful identities, policy evidence, and event links without retaining sensitive prompts or outputs unnecessarily.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Know what hashes, signatures, and anchors can establish

Integrity mechanisms support particular claims, not the whole authorization case. A hash can help detect changes to data; a signature can identify a signer; an external anchor can support a claim that evidence existed at a certain time. None alone proves that the recorded content is semantically true, that the signer had authority to approve the action, or that every relevant event was captured.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The September 8, 2026 evidence-model working paper separates artifact integrity, temporal existence, provenance, approval evidence, declared ordering, capture claims, relevance, monitoring, and policy assessment. It also cautions that its conceptual model does not validate a particular implementation, prevent every failure, or automate legal compliance. Treat these properties as separate checks rather than using “tamper-evident” as shorthand for “authorized and complete.”

Read protocol and prototype claims at the right level

Microsoft’s Agent Governance Toolkit protocol is project documentation and architecture material; its design does not show that a particular deployment implements or enforces it. The IETF text is an Internet-Draft, not a binding finalized standard. The evidence-model paper and the April 2026 paper Proof of Execution: Runtime Verification for Governed AI Agent Actions are preprints or working papers, not certifications or independent production validations.

The Proof of Execution authors report approximately 2.7 ms of overhead for a minimal single-node TypeScript flow and 4.4% for concurrent batch workloads. These are measurements reported for their prototype and conditions, not general performance guarantees. The paper’s proposed architecture separates planning, enforcement, effect, and recordkeeping, and describes an attestation object connecting authorization, enforcement, durable effect, tamper-evident history, and replay context.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.