Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstalliTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
LLM agent costs are hard to attribute because a user-visible task is a workflow, but token usage accrues request by request. One task may trigger multiple model calls, tool results fed back into the model, retries, handoffs, and delegated agents. To make the numbers useful, record usage for each model request, connect each event to explicit run and agent boundaries, then roll it up without counting nested work twice. Keep provider-reported usage, estimates, and unknown values distinct.
Why one agent task can create many token charges
A user sees one task, but an agent may make several model calls before it finishes. Each request can include instructions, tool definitions, conversation history, user input, files or images, and results returned by tools. A response may contain ordinary text, tool-call arguments, or reasoning tokens; OpenAI documents that reasoning tokens are billed as output tokens. Its guidance therefore recommends summing usage across the calls made to complete a task, rather than treating the first or final request as the whole task. OpenAI’s agent usage guidance
Tools add another boundary. A client-executed tool can return data that becomes input to a later model request, increasing token usage indirectly. A retry can issue another request, while a handoff can cause a different agent to continue the work. Tool execution may also have costs of its own, including provider-hosted tools, sandbox compute, or third-party services. Those costs should be tracked separately from token usage rather than folded into a token total.
Which boundary should a cost report use?
Different boundaries answer different questions. A request-level record helps diagnose a spike; a run-level total answers what a particular task used. Agent-level records show where delegated work occurred, while customer or feature rollups support ownership and budgeting. Preserve the underlying request events so each higher-level figure can be explained.
#1 Best Overall
| Boundary | What it answers | What to watch for |
|---|---|---|
| Model request or generation | Which request used the tokens, and which model or provider handled it? | One task commonly contains multiple requests; a request total alone is not a task total. |
| Agent invocation | How much usage belongs to this root agent or delegated agent invocation? | Do not assume a parent agent’s aggregate includes or excludes its children without checking the framework’s definition. |
| Run or workflow | What usage was associated with the user-visible task, including retries and delegated work? | Roll up each provider request once, even when multiple spans describe the workflow. |
| Customer, team, or feature | Which known owner or product area should receive the run’s usage? | Ownership requires explicit context; it cannot reliably be inferred from token counts. |
The OpenAI Agents SDK aggregates usage across model calls in a run, including calls that produce tool calls or handoffs, and exposes per-request usage entries for more detailed inspection. Those levels serve different purposes: the aggregate is convenient for totals, while request entries help locate which call produced a large input or output footprint. OpenAI Agents SDK usage documentation
How nested agents and retries distort naive totals
Delegation creates a tree of work, not a single flat sequence. OpenAI’s tracing guide says an agent span’s usage covers that agent alone and excludes its subagents. OpenTelemetry likewise recommends invocation-scoped inference and tool-call counts: activity performed by a child agent belongs to that child’s invocation. OpenAI tracing guidance OpenTelemetry GenAI metric conventions
Framework totals can follow different rules, so a rollup must document whether a run-level aggregate already includes child-agent calls. The safe accounting rule is to retain one event per provider request and count that event once in the chosen run total. Treat delegation and handoff as relationships between events, not as copies of the child’s usage. Retries also need their own request records: they consume additional model work even when an application ultimately returns only one answer.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
What a useful request-level usage record contains
Capture detail before aggregating. At minimum, associate each model request with its provider and model, a request or generation identifier when available, its run and agent invocation, and the usage fields returned by the provider. Preserve the original provider payload when the adapter supports it; normalized fields alone may omit provider-specific billing detail.
- Input, output, and total token counts when supplied, preserving whether each value came directly from the provider or was derived.
- Cache-read and cache-write usage when exposed, with their relationship to total input made clear.
- Reasoning-token detail when available; for OpenAI usage, reasoning tokens are included in output tokens.
- Provider-reported billed units when they differ from model-consumed token counts. OpenTelemetry recommends reporting billed units in that case so telemetry reflects the units charged. OpenTelemetry GenAI span conventions
- A usage status such as provider-reported, derived, estimated, pending, or unknown, plus the time the record was observed or updated.
Cached input is part of total input, while cache-read and cache-creation figures, when present, are details about subsets of usage. Do not add a subset to its parent total as if it were extra input. OpenAI’s Agents API usage fields do not separately expose cache-write counts, so those fields may not be enough to calculate exact charges when a pricing model bills cache writes separately. OpenAI’s agent usage guidance
Adapters can affect what survives into telemetry. The Agents SDK warns that some provider adapters require an explicit usage option, and normalized usage may lose provider-specific information unless raw usage preservation is supported and enabled. Check the actual provider, adapter, and streaming configuration deployed; retaining a raw snapshot cannot recover usage a provider never returned.
Rank #3
How to structure traces and rollups
Use a trace tree to represent causality, then compute totals from the request-level usage events attached to that tree. A practical structure has one workflow or run boundary for the user-visible job, agent spans for the root and delegated workers, a generation record for each model request, and tool spans for tools executed by the application. Propagate a stable run identifier and parent-child context through each delegation, generation, and client-side tool call. OpenAI’s trace model provides sessions, turns, agent spans, generation spans, and tool spans, and distinguishes root agents from subagents. OpenAI tracing guidance
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →- Record each request once. Store the request-level usage event with its provider/model identity and run, agent, and generation context.
- Connect, do not duplicate. Link child-agent and tool activity to the parent workflow through trace relationships or identifiers. Do not copy child usage into a parent event.
- Define each rollup. Publish totals at request, agent invocation, run, and known owner boundaries. State whether framework-provided aggregates include nested agents.
- Reconcile totals to their inputs. Keep raw events available so a run total can be recomputed and traced back to individual requests, including failed or retried calls.
This design is a practical synthesis of OpenTelemetry’s invocation-scoped metrics guidance and framework-specific aggregation behavior; it is not a guarantee that every SDK reports identical totals. OpenTelemetry’s span conventions also encourage developers to instrument tools invoked by their own code when automatic instrumentation does not cover them. OpenTelemetry GenAI span conventions
How to keep token usage, estimated cost, and final charges distinct
Token usage is an input to cost accounting, not always a complete bill. Store counts separately from estimated currency amounts. If estimating cost, apply a versioned provider-and-model price table and retain the price-table version and calculation time alongside the result. Add provider-specific billable categories only when their definitions are documented. A calculated estimate is not an invoice amount: pricing dimensions can change, and usage fields may omit categories that affect charges.
Missing usage is not zero usage. OpenAI documents that trace usage can be null when unknown, may arrive after an agent turn ends, and can change as it becomes available; trace usage is best-effort and is not necessarily a final bill. Represent such records as pending or unknown and update them if more complete usage arrives. OpenAI tracing guidance
Provider conventions also matter. OpenTelemetry advises reporting billed token units when a provider distinguishes them from model-consumed tokens. Its attribute registry defines common GenAI usage concepts, but the conventions are evolving and should not be mistaken for a universal guarantee that every provider or instrumentation exposes every field. OpenTelemetry GenAI attribute registry
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why pair traces with metrics?
Traces explain an individual run: which agent acted, what it called, where it handed off work, and which request contributed usage. Metrics show aggregate trends without requiring a person to inspect every trace. OpenTelemetry’s GenAI metric conventions recommend invocation-scoped inference and tool-call measurements, include failed client-side operations, and assign delegated work to the child invocation so the tree can be counted once. OpenTelemetry GenAI metric conventions
Keep high-cardinality identifiers, such as individual run IDs, in traces rather than using them as metric dimensions. Also document the coverage boundary: OpenTelemetry’s tool-call metric covers client-side calls, not tools executed inside a model provider’s service, such as provider-hosted web search or code execution. Represent those provider-side operations separately if they matter to your cost or operational accounting.
What telemetry cannot establish on its own
Instrumentation quality depends on the full deployed path: provider, framework, adapter, streaming mode, and tool execution model. A normalized record can be incomplete; a trace can be delayed; and a token count may not include every chargeable category. Validate what is actually emitted before treating a dashboard as a financial source of truth.
Tracing also has data-governance implications. OpenAI notes that traces may contain prompts, tool arguments, and results. Its guide says trace export requires organization trace export to be enabled and appropriate project API-key permissions; export is not automatically enabled for future delivery. Set retention and redaction rules to match your data and security requirements, since the cited documentation does not prescribe a universal policy.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
A practical implementation checklist
- Assign a stable identifier to each user-visible workflow and propagate it to root agents, subagents, model requests, and client-side tools.
- Capture one usage event for every model request, including retries and calls that lead to a handoff.
- Retain provider/model identity, available request identifiers, returned token classes, usage provenance, and raw usage payloads where supported.
- Define whether each SDK or framework aggregate includes delegated-agent usage before combining its totals with span-level records.
- Count provider requests once in run and owner rollups; keep tool and non-token costs in distinct categories.
- Represent absent or late usage as unknown or pending, not zero, and distinguish estimates from provider-reported values.
- Use traces for causal diagnosis and metrics for aggregate trends; document the treatment of provider-hosted tools.
- Validate trace export permissions and establish data handling policies before capturing prompts or tool content.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

