Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To make an AI agent observable, instrument each run as a trace with nested spans for model calls, tools, handoffs, and important application steps. Add structured logs for discrete events and errors, plus metrics for trends such as run volume, failures, and latency. A framework’s tracing SDK can help capture the execution path, but it does not automatically provide a complete logs-and-metrics strategy.

What to instrument in an agent run

Start a trace at the boundary of a coherent task, such as a user request or background job. Give it a stable run identifier and only the workflow metadata needed to find and compare executions. Nest spans beneath that trace so the record shows the path an agent took, not merely a flat list of events.

  • Model generations: capture timing, outcome, and any permitted input or output data.
  • Tool calls: record which operation ran, whether it succeeded, and how long it took; treat arguments and results as potentially sensitive.
  • Handoffs and guardrails: mark changes in control flow, validation, or policy decisions.
  • Application work: add spans around retrieval, external services, retries, and custom decision boundaries when they help explain latency, errors, or surprising results.

The OpenAI Agents SDK includes built-in tracing for agent-run events and nests spans under the current span. Its documented event coverage includes LLM generations, tool calls, handoffs, guardrails, and custom events. See the OpenAI Agents SDK tracing guide. With OpenTelemetry-based instrumentation, use the supported instrumentation for your framework and add manual spans for important work it does not cover.

Add logs and metrics for different questions

A trace explains one execution. Logs record discrete events that operators need to search, such as a retry, a state transition, or an error. Make logs structured and include a trace or run identifier where possible, so an individual event can lead back to its execution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Metrics answer aggregate operational questions. Useful starting points include:

  • Run volume and completion or failure rate.
  • End-to-end and operation-level latency.
  • Retry frequency.
  • Resource or token usage, when the source provides it.

Choose metric names and dimensions supported by your instrumentation and backend. Avoid high-cardinality labels—such as unique user or run IDs—and do not put secrets or raw prompt content in metric dimensions. The official guidance cited here gives tracing and usage examples, but does not establish a universal agent-metrics schema; define metrics around the service questions you need to answer.

Choose an instrumentation and export path

Instrumentation and destination are separate decisions. A framework SDK may collect useful spans while allowing different processors or exporters; OpenTelemetry provides a portable route to supported destinations. Compare the options against the framework and language you actually run, automatic versus manual coverage, export format, hosting and access controls, retention and deletion behavior, and the effort and overhead required to operate them.

Path What the documentation supports What to verify
OpenAI Agents SDK tracing Built-in tracing and configurable trace processors; traces can be routed through custom processors. OpenAI SDK documentation Current SDK release behavior and whether replacing default processors also removes the default OpenAI exporter.
OpenAI Agents API trace workflow Dashboard inspection and session-trace export as OTLP JSON through the documented API. OpenAI API documentation Organization export enablement, project API-key permissions, pagination, and whether a manual export workflow meets delivery needs.
Google Cloud with OpenTelemetry Google recommends OpenTelemetry and provides LangGraph and ADK samples. Google Cloud documentation Framework and language coverage, deployment-specific setup, access controls, and storage and retention policies.
Amazon CloudWatch with OpenTelemetry AWS documents Python and Node.js paths for LangGraph, LangChain, Strands Agents, CrewAI, OpenAI Agents, LlamaIndex, and Vercel AI SDK. AWS documentation Prerequisites and routing for the specific framework and compute environment; support listed by another provider does not establish AWS support.

Confirm coverage and behavior for the versions, regions, policies, and deployment environment you use. A one-time export is not the same as automatic delivery of future traces.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect and export OpenAI agent traces

OpenAI Agents SDK

The SDK traces the runner by default and supports trace processors. If you replace the default processors, check whether the default OpenAI exporter remains active; custom routing can change where data is sent. The SDK’s tracing page documents processor behavior and sensitive-data controls at openai-agents-python/tracing.

OpenAI Agents API

  1. Inspect: in the dashboard, open Logs → Agents, then follow the session → turn → step hierarchy to examine completed work.
  2. Enable and authorize export: export must be enabled for the organization, and the caller needs suitable project API-key permissions.
  3. Request the session trace: use /v1/agents/sessions/{session_id}/traces to export traces as OTLP JSON. The endpoint is paginated, so handle all pages if you need the full session record.

These inspection and export details are documented in the OpenAI Agents API tracing guide. Its example session records 252,468 tokens; that is an illustrative session in the documentation, not a typical-run figure or benchmark. The guide also notes that usage may arrive after a turn, may be null when unknown, and can change, so it should not automatically be treated as a final bill.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Protect payloads and plan for incomplete telemetry

Prompts, completions, tool inputs, tool outputs, and audio can contain sensitive data. Decide explicitly which, if any, belong in telemetry. In the OpenAI Agents Python SDK documentation, trace_include_sensitive_data defaults to true, and generation and function spans can store inputs and outputs. Review the current SDK controls before production and disable capture where it is not needed.

If redaction must happen before delivery, do not assume that adding a redaction processor alongside the default exporter guarantees that the exporter sees only redacted data. The SDK documentation describes replacing processors and placing redaction and delivery in an application-owned exporter. Use allowlists for captured fields, minimize identifiers, restrict access to exported data, and ensure exporter failures do not print payloads.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google Cloud recommends storing prompts and responses in Cloud Storage rather than log entries, allowing finer-grained control such as deleting an individual stored conversation. Google’s page, last updated 2026-09-30 UTC, states that Cloud Logging’s maximum log-entry size is 256 KiB; an oversized entry can be rejected, and fields over their limits may be truncated. Individual Cloud Logging entries cannot be deleted. Design consumers and alerts to tolerate missing or truncated spans, and decide how retries and export failures should be handled. See Google Cloud’s agent observability guidance.

Roll out with a practical checklist

  • Trace one complete run, from request or job entry through its final outcome.
  • Check that model, tool, handoff, retry, and relevant custom-work spans retain parent-child relationships.
  • Emit structured logs for state changes and errors, correlated with a run or trace identifier.
  • Define a small set of operational metrics and keep sensitive or high-cardinality values out of dimensions.
  • Verify exporter destination, permissions, pagination or continuous-delivery behavior, and failure handling.
  • Test capture and redaction settings with representative sensitive inputs before enabling production telemetry.
  • Set access, retention, and deletion rules for traces, logs, metrics, and any separately stored payloads.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.