Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A trace can show that an AI agent called a model, ran a tool, and stopped. None of those facts alone proves that the requested work is complete or that its result reached the person or system expecting it. For long-running agent work, the useful question is not simply whether execution ended: it is whether a defined deliverable was verified where its consumer should see it.

There is no universal, finalized status vocabulary for that kind of agent task established by the standards and guidance discussed here. Existing proposals and adjacent conventions offer useful pieces, but they serve different purposes.

What does “done” mean for an AI agent?

Several events that dashboards may color green describe different things:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • A span ended: one recorded operation, such as a model request or tool invocation, stopped. It says nothing by itself about the whole task.
  • A tool call succeeded: a tool returned successfully. The agent may still need to interpret the result, take more steps, or deliver an output.
  • The agent run reached a terminal state: execution stopped, perhaps because it succeeded, failed, timed out, or was cancelled. A terminal state is not necessarily success.
  • The deliverable was verified for its consumer: the expected result is visible on the surface where the person or system should receive it. This is the strongest basis for reporting the requested work as done.

Collapsing these into one success indicator hides where a task actually stopped. A completed tool call can coexist with an incomplete user request; an ended run can be a failure; and a generated answer can exist internally without reaching its intended destination.

What status vocabularies already exist?

OpenTelemetry CI/CD conventions

OpenTelemetry’s CI/CD semantic conventions define task result values including success, failure, error, skip, cancellation, and timeout, as well as pipeline states pending, executing, and finalizing. The conventions are labeled Release Candidate and apply to CI/CD; they are not proof of a universal agent-task standard. OpenTelemetry CI/CD semantic conventions.

Agent Arc Status Protocol draft

The Agent Arc Status Protocol v0.2 draft proposes phases named started, milestone, heartbeat, done, and blocked for long-running, authorized agent work. Its key distinction is that an emitter must verify completion from the consumer’s vantage point before sending done—the deliverable must be visible on the surface the consumer expects. Reporting an incomplete task as done violates the draft’s conformance rule. The protocol is a draft, not a finalized standard. Agent Arc Status Protocol v0.2.

Agent Runtime Telemetry System draft

An IETF Internet-Draft dated July 2026 describes a broader telemetry framework that includes task completion and output-validation signals. It remains a working document, not a finalized IETF standard; its stated expiration date is January 7, 2027. IETF Agent Runtime Telemetry System draft.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpAMP is for telemetry-agent management

OpenTelemetry’s Open Agent Management Protocol (OpAMP) handles management and status reporting for telemetry collection agents, including outcomes of package installation. It does not define whether an AI assistant has finished a user’s request. OpAMP is marked Beta. OpenTelemetry OpAMP specification.

What should agent telemetry record at completion?

Keep detailed execution tracing separate from the task’s lifecycle outcome. OpenTelemetry spans can show how reasoning, model, tool, memory, retrieval, and handoff operations unfolded; a task-level event or metric can say whether the requested work reached a verified outcome. AWS recommends spans across those operations and custom metrics such as task success and failure rates when they are not captured implicitly. That is vendor implementation guidance, not an open standard. AWS guidance for generative AI observability in CloudWatch.

A practical task-level record should identify the work, show how its state changed, and make its final outcome auditable. The following is a design checklist, not a standardized field schema:

  • Stable task identifier: correlate progress and outcome events for the same unit of work.
  • Start and update timestamps: show when work began and when status last changed.
  • Milestones or heartbeats: distinguish active progress from a task that has gone silent.
  • Explicit blocked and failure states: make non-success outcomes visible rather than treating every stopped run as done.
  • Terminal outcome: record success, failure, cancellation, timeout, or another clearly defined ending appropriate to the system.
  • Consumer-facing completion check: state what surface or delivery condition was checked before marking the task complete.

For a long-running task, the Agent Arc draft gives a default cadence floor of five minutes and a default silence window of twenty minutes. Those are defaults in that draft, not universal operating requirements. Teams should treat cadence and silence as explicit policy choices rather than assuming a heartbeat proves progress or completion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How can teams implement this without confusing the signals?

  1. Define the deliverable and consumer surface. Before work starts, specify what result counts as complete and where it must appear—for example, a report in a designated destination or a change visible in the target system. Without this, “verified” has no consistent meaning.
  2. Assign a task identifier and emit lifecycle updates. Carry the identifier across asynchronous boundaries so progress events can be joined to execution traces. Emit a start event, meaningful milestones or heartbeats, any blocked state, and a terminal outcome.
  3. Trace execution separately. Use spans to inspect model, tool, memory, retrieval, and handoff operations. Treat a successful span or tool response as evidence about that operation, not as proof of the whole task’s completion.
  4. Verify delivery before emitting success. Check the consumer-facing destination or state, then mark the task done only if the expected deliverable is visible there. If the check fails or cannot be made, use an appropriate non-success or unresolved state instead of a green success.
  5. Choose instrumentation based on the platform and the task. Built-in vendor instrumentation can reduce setup effort on supported platforms; framework-specific spans plus custom task events or metrics may offer more portability or better task-level coverage, at the cost of implementation and operational work. Compare options by whether they capture user-task outcomes, correlate across agent and tool boundaries, verify consumer-visible delivery, and fit the required portability and cost—not by assuming feature parity.

AWS documents one implementation path through CloudWatch tracing and custom metrics. The Agent Arc draft offers a transport-agnostic status vocabulary. Neither makes the other interchangeable, and neither establishes a benchmark between approaches. Use detailed tracing to explain how the work ran; use a verified task outcome to answer whether it was actually completed.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.