Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

LLM agent costs are hard to attribute because a user-visible task is a workflow, but token usage accrues request by request. One task may trigger multiple model calls, tool results fed back into the model, retries, handoffs, and delegated agents. To make the numbers useful, record usage for each model request, connect each event to explicit run and agent boundaries, then roll it up without counting nested work twice. Keep provider-reported usage, estimates, and unknown values distinct.

Why one agent task can create many token charges

A user sees one task, but an agent may make several model calls before it finishes. Each request can include instructions, tool definitions, conversation history, user input, files or images, and results returned by tools. A response may contain ordinary text, tool-call arguments, or reasoning tokens; OpenAI documents that reasoning tokens are billed as output tokens. Its guidance therefore recommends summing usage across the calls made to complete a task, rather than treating the first or final request as the whole task. OpenAI’s agent usage guidance

Tools add another boundary. A client-executed tool can return data that becomes input to a later model request, increasing token usage indirectly. A retry can issue another request, while a handoff can cause a different agent to continue the work. Tool execution may also have costs of its own, including provider-hosted tools, sandbox compute, or third-party services. Those costs should be tracked separately from token usage rather than folded into a token total.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which boundary should a cost report use?

Different boundaries answer different questions. A request-level record helps diagnose a spike; a run-level total answers what a particular task used. Agent-level records show where delegated work occurred, while customer or feature rollups support ownership and budgeting. Preserve the underlying request events so each higher-level figure can be explained.

Boundary What it answers What to watch for
Model request or generation Which request used the tokens, and which model or provider handled it? One task commonly contains multiple requests; a request total alone is not a task total.
Agent invocation How much usage belongs to this root agent or delegated agent invocation? Do not assume a parent agent’s aggregate includes or excludes its children without checking the framework’s definition.
Run or workflow What usage was associated with the user-visible task, including retries and delegated work? Roll up each provider request once, even when multiple spans describe the workflow.
Customer, team, or feature Which known owner or product area should receive the run’s usage? Ownership requires explicit context; it cannot reliably be inferred from token counts.

The OpenAI Agents SDK aggregates usage across model calls in a run, including calls that produce tool calls or handoffs, and exposes per-request usage entries for more detailed inspection. Those levels serve different purposes: the aggregate is convenient for totals, while request entries help locate which call produced a large input or output footprint. OpenAI Agents SDK usage documentation

How nested agents and retries distort naive totals

Delegation creates a tree of work, not a single flat sequence. OpenAI’s tracing guide says an agent span’s usage covers that agent alone and excludes its subagents. OpenTelemetry likewise recommends invocation-scoped inference and tool-call counts: activity performed by a child agent belongs to that child’s invocation. OpenAI tracing guidance OpenTelemetry GenAI metric conventions

Framework totals can follow different rules, so a rollup must document whether a run-level aggregate already includes child-agent calls. The safe accounting rule is to retain one event per provider request and count that event once in the chosen run total. Treat delegation and handoff as relationships between events, not as copies of the child’s usage. Retries also need their own request records: they consume additional model work even when an application ultimately returns only one answer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a useful request-level usage record contains

Capture detail before aggregating. At minimum, associate each model request with its provider and model, a request or generation identifier when available, its run and agent invocation, and the usage fields returned by the provider. Preserve the original provider payload when the adapter supports it; normalized fields alone may omit provider-specific billing detail.

  • Input, output, and total token counts when supplied, preserving whether each value came directly from the provider or was derived.
  • Cache-read and cache-write usage when exposed, with their relationship to total input made clear.
  • Reasoning-token detail when available; for OpenAI usage, reasoning tokens are included in output tokens.
  • Provider-reported billed units when they differ from model-consumed token counts. OpenTelemetry recommends reporting billed units in that case so telemetry reflects the units charged. OpenTelemetry GenAI span conventions
  • A usage status such as provider-reported, derived, estimated, pending, or unknown, plus the time the record was observed or updated.

Cached input is part of total input, while cache-read and cache-creation figures, when present, are details about subsets of usage. Do not add a subset to its parent total as if it were extra input. OpenAI’s Agents API usage fields do not separately expose cache-write counts, so those fields may not be enough to calculate exact charges when a pricing model bills cache writes separately. OpenAI’s agent usage guidance

Adapters can affect what survives into telemetry. The Agents SDK warns that some provider adapters require an explicit usage option, and normalized usage may lose provider-specific information unless raw usage preservation is supported and enabled. Check the actual provider, adapter, and streaming configuration deployed; retaining a raw snapshot cannot recover usage a provider never returned.

How to structure traces and rollups

Use a trace tree to represent causality, then compute totals from the request-level usage events attached to that tree. A practical structure has one workflow or run boundary for the user-visible job, agent spans for the root and delegated workers, a generation record for each model request, and tool spans for tools executed by the application. Propagate a stable run identifier and parent-child context through each delegation, generation, and client-side tool call. OpenAI’s trace model provides sessions, turns, agent spans, generation spans, and tool spans, and distinguishes root agents from subagents. OpenAI tracing guidance

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Record each request once. Store the request-level usage event with its provider/model identity and run, agent, and generation context.
  2. Connect, do not duplicate. Link child-agent and tool activity to the parent workflow through trace relationships or identifiers. Do not copy child usage into a parent event.
  3. Define each rollup. Publish totals at request, agent invocation, run, and known owner boundaries. State whether framework-provided aggregates include nested agents.
  4. Reconcile totals to their inputs. Keep raw events available so a run total can be recomputed and traced back to individual requests, including failed or retried calls.

This design is a practical synthesis of OpenTelemetry’s invocation-scoped metrics guidance and framework-specific aggregation behavior; it is not a guarantee that every SDK reports identical totals. OpenTelemetry’s span conventions also encourage developers to instrument tools invoked by their own code when automatic instrumentation does not cover them. OpenTelemetry GenAI span conventions

How to keep token usage, estimated cost, and final charges distinct

Token usage is an input to cost accounting, not always a complete bill. Store counts separately from estimated currency amounts. If estimating cost, apply a versioned provider-and-model price table and retain the price-table version and calculation time alongside the result. Add provider-specific billable categories only when their definitions are documented. A calculated estimate is not an invoice amount: pricing dimensions can change, and usage fields may omit categories that affect charges.

Missing usage is not zero usage. OpenAI documents that trace usage can be null when unknown, may arrive after an agent turn ends, and can change as it becomes available; trace usage is best-effort and is not necessarily a final bill. Represent such records as pending or unknown and update them if more complete usage arrives. OpenAI tracing guidance

Provider conventions also matter. OpenTelemetry advises reporting billed token units when a provider distinguishes them from model-consumed tokens. Its attribute registry defines common GenAI usage concepts, but the conventions are evolving and should not be mistaken for a universal guarantee that every provider or instrumentation exposes every field. OpenTelemetry GenAI attribute registry

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why pair traces with metrics?

Traces explain an individual run: which agent acted, what it called, where it handed off work, and which request contributed usage. Metrics show aggregate trends without requiring a person to inspect every trace. OpenTelemetry’s GenAI metric conventions recommend invocation-scoped inference and tool-call measurements, include failed client-side operations, and assign delegated work to the child invocation so the tree can be counted once. OpenTelemetry GenAI metric conventions

Keep high-cardinality identifiers, such as individual run IDs, in traces rather than using them as metric dimensions. Also document the coverage boundary: OpenTelemetry’s tool-call metric covers client-side calls, not tools executed inside a model provider’s service, such as provider-hosted web search or code execution. Represent those provider-side operations separately if they matter to your cost or operational accounting.

What telemetry cannot establish on its own

Instrumentation quality depends on the full deployed path: provider, framework, adapter, streaming mode, and tool execution model. A normalized record can be incomplete; a trace can be delayed; and a token count may not include every chargeable category. Validate what is actually emitted before treating a dashboard as a financial source of truth.

Tracing also has data-governance implications. OpenAI notes that traces may contain prompts, tool arguments, and results. Its guide says trace export requires organization trace export to be enabled and appropriate project API-key permissions; export is not automatically enabled for future delivery. Set retention and redaction rules to match your data and security requirements, since the cited documentation does not prescribe a universal policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical implementation checklist

  • Assign a stable identifier to each user-visible workflow and propagate it to root agents, subagents, model requests, and client-side tools.
  • Capture one usage event for every model request, including retries and calls that lead to a handoff.
  • Retain provider/model identity, available request identifiers, returned token classes, usage provenance, and raw usage payloads where supported.
  • Define whether each SDK or framework aggregate includes delegated-agent usage before combining its totals with span-level records.
  • Count provider requests once in run and owner rollups; keep tool and non-token costs in distinct categories.
  • Represent absent or late usage as unknown or pending, not zero, and distinguish estimates from provider-reported values.
  • Use traces for causal diagnosis and metrics for aggregate trends; document the treatment of provider-hosted tools.
  • Validate trace export permissions and establish data handling policies before capturing prompts or tool content.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.