Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteTo debug an AI agent, trace the whole workflow—not just its final model response. A useful trace groups the run and records nested steps such as model generations, tool calls, handoffs, retrieval, and application work, with timing and outcome details. Logs make events searchable; traces show how related operations fit together. Neither alone proves that an answer is accurate or safe.
What observability means for an AI agent
A user may experience one task, while the agent performs several operations to complete it. A trace makes those operations inspectable as a connected workflow. The OpenAI Agents SDK documents default tracing for model generations, function or tool calls, handoffs, guardrails, and custom events; AWS OpenSearch documentation describes hierarchical traces across orchestration, model calls, tools, and retrieval.
Use the following as a practical mental model, not a universal naming scheme. Frameworks differ in how they represent runs, sessions, and turns.
- Trace: the record that groups an end-to-end workflow or operation.
- Span: a record of one operation, with start and end timing, status, and any captured attributes or content.
- Parent and child spans: the nesting that shows which operations happened within a larger operation, such as a model call or tool action.
- Session and turn: in the OpenAI Agents API terminology, a session can contain several turns, and a turn’s trace can group steps such as model responses, tool calls, and delegated work.
Structured logs and traces answer different questions. Logs are useful for searching individual events and application context; traces show how related operations fit together, where time accumulated, and where an error occurred. A final answer is not an execution trace: it can hide multiple generations, retries, tools, handoffs, guardrails, and retrieval steps.
Recommended Free Tools
#1 Best Overall
What to instrument
Record the full execution path your team controls, especially operations that can change the result, fail, or add significant latency. Start with the workflow invocation and add relevant child spans for:
- Each model generation, including provider and model identifiers when available.
- Tool execution, with the tool name and call identifier, and arguments or results only when appropriate to capture.
- Delegation or handoff between agents.
- Retrieval activity, such as the search or lookup operation that supplies context.
- Guardrails and application-specific work that materially affects the outcome but is not already represented.
Choose stable identifiers and useful, low-cardinality dimensions for filtering and grouping. OpenTelemetry’s GenAI conventions provide guidance for workflow names and conversation IDs. They advise against inventing a conversation ID when one is unavailable: do not substitute a random UUID, trace ID, or hash of request content. Use a conversation ID only when the instrumented library already has one or the application supplies it.
Rank #2
Do not assume automatic instrumentation captures every internal step. Coverage depends on the specific library and configuration. Inspect an exported trace from a representative run; add custom spans where meaningful work is missing. OpenTelemetry conventions are a living project document, so check the current guidance when implementing them.
How to investigate a failed, incorrect-looking, or slow run
- Find the run. Filter using identifiers your application records, then locate the relevant run or turn and time window. The OpenAI Agents API trace UI documents filtering by model, status, or date and opening a session timeline.
- Follow the trace tree and timeline. Start at the workflow or agent root, then inspect child spans for model responses, tools, retrieval, and delegated work. Look for the first failed span, unexpected result, retry, or unusually long operation. The timeline helps show ordering, overlap, duration, and outcome status.
- Inspect the span details. When content capture is enabled, compare model inputs and outputs or tool arguments and results. Check provider and model identifiers, tool name and call ID, status, error, duration, and token usage where available. A blank or unknown usage value is not necessarily zero: the OpenAI guide notes that usage may arrive after a turn and can change as it becomes available.
- Reproduce or isolate the suspect operation. Use the trace to identify the operation and its surrounding context, then reproduce it with appropriately sanitized inputs or test the tool/model boundary independently. A trace narrows the investigation; it is not a universal incident-response procedure.
- Close instrumentation blind spots. Add custom spans for important application work not already represented, using names and attributes that support filtering. SDKs may provide custom span and processor mechanisms.
A trace records what the instrumented system observed: operations, captured inputs or outputs, status, errors, and timing. It can help localize an execution fault or explain latency, but it does not certify factual quality, policy compliance, or safety. Those judgments require appropriate evaluation and review beyond execution telemetry.
Rank #3
Choosing built-in tracing or OpenTelemetry
Two documented approaches are to use tracing built into an agent SDK or instrument the workflow with OpenTelemetry and send it to a compatible backend. They are not mutually exclusive in every architecture; choose based on the libraries and operational environment you actually use, then verify the exported structure.
| Approach | What it offers | What to verify |
|---|---|---|
| Framework or SDK built-in tracing | The OpenAI Agents SDK documents automatic traces and spans, runtime defaults, sensitive-data settings, and export processors. This is a direct starting point for applications using that SDK. | Defaults depend on implementation and runtime. The JavaScript SDK documentation says tracing is enabled by default in server runtimes and disabled by default in browsers and test mode; the Python documentation describes tracing as enabled by default. Check the exact package version, runtime, and configuration in use. |
| OpenTelemetry instrumentation plus a backend | OpenTelemetry GenAI conventions provide shared naming and attribute guidance. AWS documents AI traces, OpenTelemetry integration, auto-instrumentation, and querying in OpenSearch, as well as manual examples for invocation and tool spans. | Confirm instrumentor coverage for each library and provider, export configuration and permissions, and the structure the backend receives. Conventions and example attributes are not a guarantee that every framework emits the same spans. |
Compare options on practical fit rather than assuming a universal winner:
- Does instrumentation cover the framework and providers you run?
- Can you see tool calls, retrieval, handoffs, and application-specific work—not only model calls?
- Does each span contain useful status, timing, and detail without collecting more content than necessary?
- Can you control sensitive-data capture and redaction?
- Can you export to the destinations you need and correlate traces with logs and metrics?
- Can the team efficiently filter, query, and investigate the resulting telemetry?
For the OpenAI Agents API, session trace export returns OTLP JSON, but export must be enabled for the organization and requires suitable project permissions. Treat export format as only one part of interoperability: confirm permissions, configuration, span coverage, and actual output.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Handle trace content as sensitive data
Depending on configuration, traces may contain prompts, model outputs, function inputs and results, or audio data. OpenAI’s JavaScript and Python Agents SDK documentation describes settings to disable sensitive-data capture; the Python documentation says its sensitive-data capture setting is enabled by default. OpenTelemetry’s GenAI conventions warn that input-message attributes can contain sensitive or personal information.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
Before enabling production capture, decide which details are genuinely needed to diagnose your system. Configure omission or redaction, restrict access to traces, and align retention with your application’s data policy. Check the configuration in the runtime you deploy rather than relying on a default that may differ across SDKs or environments.
Quick Recap
Documentation for implementation details
- OpenAI Agents SDK for JavaScript: Tracing
- OpenAI Agents SDK for Python: Tracing
- OpenAI Agents API: Agent traces
- Amazon OpenSearch Service: Generative AI traces
- OpenTelemetry GenAI agent span conventions
- OpenSearch: Manual OpenTelemetry instrumentation
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

