Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Your agent’s cost problem may not be the model alone. To find out, measure each model request, tool call, retrieval, retry, and handoff in a completed task, then compare the total with the task’s outcome. The claim that unmeasured steps are the main cost driver is a hypothesis—not a finding established across workloads. Your own traces and billing records can show whether it is true for your agent.

What to measure in an agent run

A provider bill or a final-response token count cannot show which parts of a multi-step workflow drove the cost. Give each user task a stable identifier, carry it through every request and workflow step, and capture enough detail to connect usage with duration and outcome.

  • Model requests: provider and model identifiers, request identifiers, input and output tokens, cached or reasoning tokens when available, and usage returned for each request.
  • Workflow steps: tool and retrieval calls, delegated-agent work, retries and attempts, start and end times, status, and duration.
  • Other charges: billable tool or API usage, hosting, sandbox compute, and third-party service costs.
  • Task context: task category, version of the workflow, and a completion result or evaluator score.

Aggregate by task category and completion outcome, not only by model call or monthly bill. Cost per successful task can be useful, but it is an accounting choice, not a universal metric prescribed by the providers.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to instrument the workflow

1. Keep a task ID across the full run

Assign each user task a stable run or task ID. Preserve it across model requests, tools, retrieval, retries, and handoffs so that separate events can be reconstructed as one workflow.

2. Record each model request and cross-check totals

Capture usage for every request, including calls that trigger a tool or handoff. The OpenAI Agents SDK reports request count, input tokens, output tokens, total tokens, and per-request usage entries; its aggregate run totals provide a useful cross-check, not a substitute for request-level records. OpenAI Agents SDK run results and usage.

A session can retain conversation history, but each run’s reported usage is independent. Earlier messages may be sent again as input on a later turn, so a run total should not be mistaken for the cumulative cost of an entire session.

3. Trace tools, retrieval, retries, and handoffs

Record each step’s type, start and end time, status, attempt or retry number, and any available billable usage. A trace can reveal work hidden behind a single final answer: model responses, tool calls, delegated agents, inputs, outputs, and duration. OpenAI tracing organizes this activity into sessions, turns, and spans. OpenAI Agents SDK tracing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Usage may be reported after a turn or remain unknown in a trace. A blank or null usage field is not proof that the step cost nothing.

4. Attach the outcome

Record whether the task completed and, where appropriate, its evaluation result or quality score. Compare cost alongside latency and quality across workflow versions; a cheaper run is not an improvement if it fails more often or produces worse work.

Why tokens are only part of the bill

Token counts help explain model charges, but they do not capture every cost or always map exactly to an invoice. Model-call inputs and outputs can include tool definitions, conversation history, tool results, and reasoning. OpenAI’s documentation states, “Reasoning tokens are billed as output tokens.” OpenAI: Observability and usage.

Estimate model charges using the applicable provider prices and the token categories actually recorded. Label the result an estimate if the usage fields omit a charge component. For example, cache-write charges may apply even when the usage fields described in the documentation do not expose a separate cache-write count. Track tool, hosting, sandbox, and third-party service charges separately rather than treating them as token costs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some observability tools calculate cost automatically for supported language-model integrations and allow manual costs for other run types. LangSmith, for example, documents both automatic cost calculation for supported integrations and manual cost assignment to tool and retrieval runs. LangSmith cost tracking. This distinction matters when an API or retrieval charge is not proportional to model tokens.

Choose tracking that covers the whole workflow

Built-in traces, third-party observability, and custom logging can all contribute to a cost ledger. The right setup is the one that captures the relevant steps and connects them to outcomes and billing—not a vendor ranking inferred from documentation examples.

Approach What the cited documentation establishes What to check for your workflow
Built-in OpenAI tracing and usage Run usage includes request and token totals with per-request entries; traces represent model responses, tools, delegated work, and recorded usage. Usage; tracing. Whether every relevant external charge and task outcome is captured in your records.
Third-party observability LangSmith documents automatic cost tracking for supported LLM integrations and manual cost assignment for other run types, including tools and retrieval. An OpenAI Cookbook example demonstrates Langfuse tracing with approximate cost based on token use and duration per step or run. LangSmith; OpenAI Cookbook Langfuse example. Integration coverage, delegated-agent and retry visibility, manual-cost support, and how estimates are reconciled with actual charges.
Custom logging A team can record its own task IDs, step events, costs, and outcome labels; the reviewed sources do not establish a universal custom schema. Whether the log preserves per-request usage and workflow detail, captures non-model charges, and can be reconciled with provider billing.

Whichever approach you use, verify that root and delegated agents are represented; model, tool, retrieval, and retry activity is covered; per-step usage and duration are available where applicable; manual non-model costs can be recorded; and runs connect to task category and outcome.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reconcile trace estimates with provider billing

Use provider-side usage and billing records to check whether your trace-based estimates align with billed charges. OpenAI’s Usage Dashboard and response usage fields offer provider reporting, but dashboard costs are not combined across separate organizations. A consistent project and account structure, or custom analysis, may be needed for consolidated reporting. OpenAI Usage Dashboard and usage guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Differences do not automatically mean the trace is wrong: check whether it includes all requests and charge categories, whether usage arrived later, and whether non-model services are billed separately. Keep estimates and invoice-reconciled amounts clearly distinguished.

Find the steps worth changing

Once the ledger joins cost, workflow steps, and outcomes, inspect high-cost and failed traces. Treat patterns as leads to investigate, not automatic proof of waste:

  • Repeated requests or retries that do not improve completion.
  • Large tool or retrieval results sent back into the model context.
  • Delegation that adds work without a measurable quality or completion benefit.
  • Expensive reasoning or output volume relative to the task result.
  • Tool, sandbox, or third-party charges that token totals do not reveal.

Test a change against comparable tasks and track cost, completion, latency, and quality together. A step is worth removing or redesigning only if the savings do not come at the expense of the result users need.

What the evidence does—and does not—show

The available documentation explains how to inspect usage and trace workflow steps; it does not quantify what share of agent costs comes from unmeasured steps or establish that those steps are usually the dominant driver. Model prices, tokenization, output volume, and reasoning can all matter. Measure the actual workload before deciding whether to change the model, the workflow, or both.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.