iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An AI agent’s bill is usually the sum of repeated model requests plus any separately metered tools, compute, storage, or services—not simply one request at the model’s advertised token rate. To understand a run’s cost, count every model call, its input and output usage, and any other billable work, including retries and delegated agents.
Why can one agent task trigger several charges?
An agent can make multiple model calls while working through a single task. OpenAI’s Agents API documentation describes this directly: “An agent may make several model calls while completing a task.” A run that plans, calls a tool, reads the result, and then answers may therefore consume tokens across several requests. If a call is retried or work is handed to a subagent, those requests and their associated usage also belong in the task’s cost.
That is why the model’s per-token rate alone cannot tell you what a completed agent task costs. You need the usage for the full workflow, not just its final answer or first request.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsWhat goes into the cost of an agent run?
A useful checklist is: model input, cached input or cache writes where applicable, model output, separately metered tools, compute and storage, third-party services, retries, and delegated work. This is an accounting framework, not a universal billing formula: providers define which categories apply and how they charge.
#1 Best Overall
| Cost area | What may be included | What to record |
|---|---|---|
| Model input | Instructions, tool definitions, conversation history, user messages, files or images, and results returned by tools. | Input tokens for each request, plus cache-related usage fields when the provider exposes them. |
| Model output | Generated text, tool-call arguments, and, for the cited OpenAI usage guidance, reasoning tokens billed as output tokens. | Output tokens for each request, including tool-call arguments where reported. |
| Tools and grounding | Some features have a separate charge; others add token usage or use the selected model’s rates. The charge depends on the provider and feature. | Tool name, number of calls or queries, and any separate usage or fee shown by the provider. |
| Compute and storage | Hosted containers, file-search storage, or other platform resources may be billed separately. | Relevant resource use and the matching platform charge. |
| Retries and delegates | Additional model requests and any tools, compute, or third-party services they use. | Which calls belong to the root run, retry, or delegated agent, where attribution is available. |
For example, OpenAI’s pricing documentation lists separately priced items such as container sessions, file-search storage, and file-search calls in addition to model-token costs. Anthropic’s pricing documentation distinguishes its web-search tool, which is charged in addition to token usage, from web fetch, for which fetched content included in model context incurs standard token costs rather than an additional web-fetch fee. Tool definitions and returned command output can also add tokens. These examples are provider- and feature-specific, not rules that apply to every tool.
Why can conversation history make later requests larger?
When a workflow carries prior messages forward, later requests may include some or all of that session history as input. OpenAI’s Agents SDK usage documentation says session history may be re-fed as input in later runs and affect later input-token counts. The effect depends on what the workflow sends: a session does not, by itself, establish how much history is included in each request.
Rank #2
Some providers can reduce the cost of eligible repeated prefixes through caching. Eligibility, cache lifetime, and billing treatment vary by model and provider. A continuing session is not proof that a request received a cache hit; check cache-usage telemetry instead of assuming a discount.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Do tool calls and retries cost extra?
Sometimes—but “tool call” does not identify a single billing method. A tool may have a separate per-call or per-query charge, consume model tokens through its definitions and results, rely on the model’s ordinary token rates, or involve billable compute or storage. A retry or delegated agent can add more model calls and may also repeat tool or infrastructure use.
For instance, Google Cloud Agent Platform pricing describes some grounding charges by query or prompt, while computer-use pricing is based on tokens sent to and generated by the model. The page specifies a January 5, 2026 billing start for certain grounding charges and July 1, 2026 effective terms for some non-global endpoints. Those dates and categories are not a substitute for checking the applicable endpoint, region, and current terms for your deployment. Provider pricing pages and implementation details can change; the documentation cited here was checked on October 7, 2026.
How do you calculate the cost of an AI agent run?
- Define the unit. Decide what counts as one task—for example, one successful user request—and identify its root run and any retries or delegated work.
- Collect every model request. For each call, record the model, input tokens, output tokens, and cache-related usage fields if supplied. Include the request’s carried-forward context, tool definitions, and tool results through the provider’s reported usage.
- Apply the matching model rates. Use the rate card for the exact model and usage categories, such as input, cached input, cache writes, or output, where those categories apply. Use the pricing terms that match your region, endpoint, and pricing date.
- Add non-token charges separately. Reconcile tool calls, grounding, hosted compute, storage, and third-party services against their own usage records and current terms. Do not assume every tool is priced per call or that every platform charge appears in token totals.
- Sum the whole workflow. Include retries and delegated work, then compare total cost per successful task as well as cost per model call. A lower-cost call does not necessarily mean a lower-cost completed outcome if it takes more attempts or fails more often.
- Reconcile with the bill. Compare run- and request-level telemetry with provider invoices or usage reports. Investigate gaps such as untracked retries, missing tool usage, or platform charges recorded outside the model’s token report.
A simple worksheet can help keep categories distinct:
Task or root run ID: __________
Model requests: __________
Input / cached input / cache writes: __________
Output: __________
Tool and grounding charges: __________
Compute and storage: __________
Third-party services: __________
Retries and delegated work: __________
Total workflow cost: __________
Successful completion? __________
What should you instrument to find the source of a high bill?
Record both run-level totals and per-request usage. OpenAI’s Agents SDK documents aggregate usage as well as request_usage_entries, which can help show which model calls consumed tokens. Preserve enough context to attribute charges across the workflow:
- Run and request identifiers, model, and whether the request belongs to the root agent, a retry, or a delegate.
- Input and output tokens, with cache-related fields when provided.
- Tool calls and their results, plus separately reported tool, grounding, compute, storage, and third-party usage.
- Whether the task completed successfully, so you can compare cost per successful task rather than only cost per call.
When tracing a spike, first identify which requests or non-token charges rose. Then inspect the contributing workflow behavior: more model calls, retries, longer carried-forward history, larger tool results, or increased separately metered activity. This turns an unexplained bill total into categories you can validate against usage records.
Best Value
How can you control costs without hiding the trade-offs?
- Measure the whole task. Compare full-workflow cost and completion rate, not a single request in isolation. The provider documentation establishes how usage can accumulate, but it does not establish a universal cost threshold at which an agent task is worthwhile.
- Use caching as a measured benefit. Where practical, keep repeated instructions and tool definitions stable, but verify actual cache usage in telemetry. Do not budget as if a cache hit is guaranteed.
- Separate token and service charges. Keep distinct records for model usage and separately billed tools or infrastructure, then confirm each feature’s current terms with its provider.
- Recheck rates and terms when budgeting. Provider prices, eligible usage categories, and effective dates can change. Match each rate to the model, feature, endpoint, region, and date that apply to the workload.
Is there a typical cost per AI agent task?
The provider rate cards and usage documentation cited here do not establish a reliable, representative cross-provider average cost per successful agent task. A useful estimate must come from a workload with its model, prompt and carried-forward context, tool mix, call count, retries, delegates, success rate, and applicable pricing terms specified. Without those details, a single “average agent cost” would conceal the factors that drive the bill.
How should you compare two agent configurations?
Compare them on the same task and account for both cost and successful completion. At minimum, include:
- Input, cached-input, and output rates for the selected models.
- Model calls required, including retries and delegated work.
- Context carried into later calls and token overhead from tool definitions and results.
- Separate tool, hosted compute, storage, grounding, and third-party charges.
- Cache eligibility and terms, alongside measured cache hits rather than assumed savings.
- Region, endpoint, pricing date, and applicable effective-date terms.
- Completion quality and success rate, so a cheaper individual call is not mistaken for a cheaper completed task.
There is no standardized comparison benchmark in the cited provider documentation. Your own instrumented workload is the meaningful basis for a cost comparison.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

