Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An agent’s “done” message is a claim, not proof. To verify the work, check whether the requested result exists in the system it was meant to change, then separately check for unwanted side effects and compliance with any consent, safety, or authorization rules.

What counts as proof that an agent finished?

Start with the user’s request and turn it into observable completion conditions. Check those conditions against the relevant system state—not just the agent’s summary, a successful tool response, or a log showing that an action was attempted.

For example, if the task was to create a calendar event, verify that the event appears on the intended calendar with the requested time and attendees. A tool log may show that the agent called an event-creation function; the resulting calendar entry is stronger evidence that the intended change occurred.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the criteria tied to what was actually requested. Microsoft Research’s 2026 work on computer-use verification warns against “phantom” rubric criteria: requirements an evaluator adds even though the user never asked for them. A check should be specific enough to assess independently without moving the goalposts.

Check the result and the process separately

A task can reach the right result through an unacceptable action, or be carried out correctly but blocked before the result is achieved. Record these as distinct judgments rather than collapsing them into a single “success” score.

  • Outcome: Did the requested change happen? Was it complete, partial, or absent?
  • Process: Did the agent follow the requested steps and avoid unwanted actions?
  • Constraints: Did it respect applicable consent, safety, and authorization requirements?

A login wall or CAPTCHA may prevent an agent from completing a task even when it handled the steps it could control appropriately. Conversely, an agent might produce the requested result while also making an unrequested change. In either case, report what happened, including whether the problem was within the agent’s control.

Include safety and consent in the acceptance rule

“The result exists” is not enough when the task has policy constraints. Before evaluating a consequential workflow, make explicit which actions require authorization, whose consent is needed, and what the agent must not do. Then count a task as acceptable only if both the requested result and those constraints are satisfied.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

IBM Research’s ST-WEBAGENTBENCH, dated July 13, 2025, pairs web-agent tasks with safety and trustworthiness policies and evaluates six dimensions. Its Completion Under Policy metric credits a task only when applicable policies are respected. In the reported evaluation of three open agents, average Completion Under Policy was less than two-thirds of nominal task completion. That finding applies to the benchmark and agents evaluated; it is not a general failure rate for agents in other settings.

How do you verify an agent’s work in practice?

  1. Write down the requested result. Convert the request into observable conditions without adding unrequested requirements. For an event, these might include the correct calendar, date, time, and invitees.
  2. Inspect the system the task was supposed to change. Look for the created or updated record, message, file, or other intended result. Prefer this state check to a self-reported status or an ambiguous success response.
  3. Check completeness and unintended changes. Confirm that all requested parts are present and look for relevant side effects. A correct record does not, by itself, show that nothing else was changed.
  4. Review the action trail when process matters. Compare the steps or tool calls with the user’s instructions and applicable constraints. Logs can help explain how an outcome occurred, but an attempted action is not proof that the resulting state is correct.
  5. Report the result with its limits. Distinguish verified completion from partial completion, a failed attempt, and a task blocked by an external condition. If a check could not be performed, say what remains unverified.

Agent-Diff, a preprint dated February 11, 2026, uses a state-diff contract: success depends on whether the expected environment change occurred. Its sandboxed evaluation covered 224 enterprise workflow tasks and nine LLMs across Slack, Box, Linear, and Google Calendar interfaces. Those benchmark details illustrate the value of checking resulting state; they are not production reliability estimates.

Can a separate verifier be trusted?

Separating verification from execution can reduce reliance on an agent’s own account of its work, but a second model is not automatically an independent source of truth. A verifier still needs clear criteria and access to suitable evidence, and it can misread both.

Microsoft Research’s April 21, 2026 article describes a Universal Verifier for web computer-use trajectories and reports results on 246 human-labeled trajectories. In that evaluation, agreement between the verifier and human labels was Cohen’s κ 0.64. The article also reports false-positive rates of at least 45% for WebVoyager and at least 22% for WebJudge relative to its human labels. These are results from that evaluation setup, not universal error rates for those tools or for agent verification generally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The same work separates rubric-based process scoring from binary outcome scoring. It reports 96 experiments in its verifier-design work and discusses the evidence-selection problem: too many screenshots can overwhelm a judge, while checking only the final screenshot can miss evidence from earlier in a trajectory. Its takeaway is practical: choose evidence relevant to each criterion and assess process separately from whether a reasonable user would consider the task done.

A public GitHub repository titled “The Verifier Agent: Mitigating Task Verification Failures in Agentic AI” proposes a Planner, Executor, and separate Verifier. The repository reports a 20-task experiment and manual ground-truth review of 120 experimental runs; the accessed page does not state a publication year. It also says its study used single-step, atomic tasks, with complex multi-step workflows left for future work. Treat those results as preliminary rather than evidence that the design has been established for longer workflows.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to ask before accepting a verification result

  • What exact user-requested conditions were checked, and were any requirements added by the evaluator?
  • Did the verifier inspect the relevant environment state, an action trail, screenshots, or only the agent’s report?
  • Were outcome, process, and policy compliance assessed separately?
  • What kinds of tasks and environments were included in the verifier’s evaluation?
  • Was it compared with human judgments, and what were the limits of that comparison?

Benchmarks answer questions about particular tasks, environments, and evaluation methods. A result from a sandbox, a set of labeled trajectories, or a small atomic-task experiment should not be read as proof that a verifier will be reliable for a different workflow.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.