Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

An agent’s tokens go to individual model requests, not to the final answer. A run-level total tells you that usage accumulated; the per-request records and the trace tree tell you where it accumulated and why. Start with the requests, then use the trace to connect each one to the tool call, retry, handoff, or nested agent that triggered it.

Why the run total is not enough

A single agent task can involve many model calls. Each call carries its own instructions, tool definitions, conversation history, tool results, and generated output. When the agent calls a tool or hands off to another agent, the work that follows is another model call, and that call repeats most of the input that came before it. The displayed answer is only the last of those calls.

The OpenAI Agents SDK aggregates usage across the whole run. Its documentation states: “Usage is aggregated across all model calls during the run, including model calls that produce tool calls or handoffs.” That aggregate is useful for budgeting, but it cannot tell you which step was responsible for a spike. For that you need the per-request entries and the trace.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cost accounting has the same structure. Input tokens, cached input, cache writes, output tokens, and reasoning tokens are reported in different categories and may be priced differently. A total that merges them hides which category is growing.

What each layer tells you

Treat three layers as different instruments, each answering a different question.

Layer Question it answers Typical source Limitation
Run or task total How much did this task consume? Aggregated usage on the run result Does not show which call caused the growth
Request-level usage Which model call consumed what, and in which category? Per-request usage entries, provider response usage fields Only covers calls the SDK or your logging captured
Trace tree What triggered each call, and in what order? Agent spans, generation spans, tool spans, handoffs Usage may be late, unknown, or absent on some spans
Organization report What did the account consume over a period? Provider usage or cost reporting APIs Aggregated across workloads; not a per-run debugger

Keep the run total, but store the child request records that produced it. A total without its children cannot be investigated later.

Field names and token categories

Provider field names differ, so a logging schema that assumes one shape will silently mis-record the other. OpenAI Chat Completions reports prompt_tokens, completion_tokens, and total_tokens. The Responses API reports input_tokens, output_tokens, and total_tokens. Map both into one internal schema, and keep the provider’s original names in the raw record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For each model request, record:

  • Provider, model identifier, and endpoint
  • Run, session, and request identifiers, plus timestamp
  • Input and output token counts, and total
  • Cached input and cache-write counts, where the provider reports them
  • Reasoning-token details, where available
  • Status and any retry relationship to an earlier request

Reasoning tokens are billed as output under the OpenAI Agents guidance, so a reasoning-heavy model can raise output cost even when its visible answer is short. The reverse also holds: OpenAI documents that reported output counts can include non-visible tokens used for formatting, tool-call structure, and message structure. A visible answer is therefore an incomplete proxy for output cost. Tokenization and output length also differ between models, so a lower per-token price does not guarantee a cheaper completed task.

Do not record a missing provider field as zero. Store it as unknown and flag it.

Reading the trace tree

Token totals tell you the size of a problem. The trace tells you its cause. Read the trace for the following:

  • Agent spans show which agent ran and how the run is structured.
  • Generation spans show the recorded input and output of each model call.
  • Tool spans show the tool called, its arguments, and its result when available.
  • Handoff and subagent spans show where control moved and which child agent consumed tokens.
  • Timing shows ordering, overlap, duration, status, and failures.

OpenAI’s tracing guide includes a worked example that shows how the numbers add up. The example records a root agent with 126,390 input and 1,567 output tokens, subagent A with 34,075 input and 465 output, and subagent B with 89,304 input and 667 output. Those child and root figures sum to 249,769 input and 2,699 output, or 252,468 total. The example’s year is not stated in the documentation. It illustrates how a trace rolls up; it is not a benchmark of typical agent usage, and it should not be read as one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Trace data can be exported as OTLP JSON. Organization-level trace export must be enabled, and the API key or project used needs the appropriate permissions.

Handling null, late, and changing usage

Usage is not always final when you read it. OpenAI says traces are built after a turn ends, and usage may arrive later, be unknown, or change. The tracing guide puts it directly: “A blank value or null means the count is unknown. It does not mean the agent used zero tokens.”

In practice this means:

  • Do not write a zero into a dashboard when a field is blank. Show the value as unknown, and count how many requests are unknown.
  • Re-read a trace before you finalize a cost figure for a run that just finished.
  • Do not treat a trace’s usage as the bill. Usage is telemetry; the invoice is computed by the provider under its current pricing.

Estimates before the request versus usage after it

Anthropic’s token counting endpoint is for planning. Its documentation describes the count as an estimate that can differ slightly from actual input usage. The preflight counter does not apply prompt caching logic, so it can overstate input for a request that would hit a cache. Some server tools are also unsupported by the counter.

Use the preflight count to decide whether a prompt fits or to project a budget. Use the usage returned with each response to record what actually happened. Do not compare the two as if they measured the same thing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Organization-level reconciliation

Anthropic’s Usage API reports organization usage in fixed time intervals. It measures uncached input, cached input, cache creation, and output tokens. Results can be filtered or grouped by API key, workspace, model, service tier, context window, residency, and speed. It also includes server-tool usage such as web search.

This view is for reconciliation and capacity questions, such as whether one workspace’s spend is climbing. It cannot tell you why a particular agent run used the tokens it did. Use it to check that your per-run records add up to the organization totals over the same period.

A diagnostic workflow

The steps below work with any provider that reports per-request usage. Each step produces a record you will need in the next.

  1. Start with a representative task and its outcome. Save the run identifier, the success or failure result, and the latency. Use comparable tasks rather than a single final response string.
  2. Expand the run into model requests. List every request, including retries, handoffs, and nested agent work. Attribute each record to its parent run and, where your system allows, to a user or workload.
  3. Separate the token categories. Keep input, cached input, cache writes, output, and reasoning details apart. Confirm the field meanings in the documentation for the specific endpoint and model before comparing across providers.
  4. Look for repeated input and loops. Check whether conversation history, tool definitions, tool output, or retries re-enter later requests. Session history is re-fed as input on later runs in the OpenAI Agents SDK, so a long-lived session can grow input on every turn even when the agent’s behavior has not changed. Whether a given agent has an actual loop is a question for its trace, not for its totals.
  5. Join usage to the trace. Inspect tool results, errors, turn order, and duration for the requests that spiked. A jump often coincides with a large tool output, a repeated call, or a longer carried history. Confirm the cause in the trace before changing the prompt or the tool.
  6. Reconcile with provider records. Use preflight counts for planning, per-request usage for runtime records, and organization reports for totals. Check current pricing before you convert tokens to dollars.
  7. Compare cost per successful outcome. Track tokens and cost per completed task beside quality and latency. A change that cuts tokens but makes the agent fail more often raises the cost of each useful result.

Common causes of token spikes

These are the patterns the trace structure makes visible. Each one should be confirmed in your own traces before you act on it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Growing history. Each request carries the earlier turns, so input climbs steadily across a long session.
  • Large tool output. A single tool result that returns a whole document or a large JSON payload is re-sent in every later request.
  • Repeated tool calls. An agent that retries a failing call produces a new model request for each attempt.
  • Verbose tool definitions. Large tool schemas are part of the input of every model call, including the ones that do not use the tools.
  • Reasoning-heavy output. Reasoning tokens bill as output, so a model set to think longer can raise cost on a short visible answer.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Comparing the tooling options

Several documented paths exist. Compare them on the dimensions below rather than choosing one as universally best.

Option Useful view Compare on
OpenAI Agents SDK usage object Run aggregation, per-request entries, cached and reasoning detail fields, session-run behavior Request granularity, how nested agent usage is aggregated, provider scope, data retention
OpenAI Agents tracing Sessions, turns, traces, root and subagent usage, tool and generation spans, OTLP JSON export Whether usage is populated when you read it, export permissions, fit with your workflow
Anthropic Usage API Organization usage by time interval, token class, and filters or groupings, including server-tool use Reconciliation at organization scope, supported dimensions, API access, integration effort
LangSmith cost tracking Automatic LLM cost from token counts and prices for documented integrations; manual costs for other run types Provider and framework coverage, custom pricing, cost attribution for non-LLM steps, data handling

Nested agents need care. The SDK documentation describes a resumed nested Agent.as_tool() run as aggregated into the active outer run, while resumed top-level checkpoints carry independent usage snapshots. Other frameworks may total nested work differently, so state which semantics your numbers use before comparing them across systems.

Measure outcomes, not just spend

Token reduction is only a win if the task still succeeds. The most useful single figure for an agent is cost per successful task, calculated from the per-request records of runs that completed. Report it alongside success rate and latency, and break it down by task type. This is a practical method rather than a published benchmark, and the right thresholds depend on your workload.

Start with the request records, follow the trace to the call that caused the growth, and only then change the prompt, history, or tool behavior. The total tells you that the agent is expensive; the requests and trace tell you what to fix.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prices and product features in this area change. Check each provider’s current pricing and documentation before turning token counts into dollars or setting a budget.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.