Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A token is a model-specific unit of content, not a word—and an AI agent’s bill can include far more than the prompt you typed and the answer you read. Instructions, conversation history, tool definitions, files, tool results, generated tool-call arguments, hidden reasoning, repeated model calls and separately priced tools may all affect the total. To estimate cost, identify the exact model and billing terms, total the usage categories reported for the full agent run, apply their current rates, and add any applicable non-token charges.

What counts as a token?

A token is a piece of content as represented by a particular model’s tokenizer. OpenAI explains that a token may represent “a character, part of a word, a whole word, or punctuation”; the count depends on the model, encoding, language and text. Its rough estimate for English—about four characters or three-quarters of a word per token—is a rule of thumb, not a dependable conversion. Spelling, capitalization, spaces and punctuation can change how text is split. See OpenAI’s token explanation.

For multimodal models, input is not limited to text. Gemini’s token guidance covers images, video and audio as well as ordinary text, and different providers may encode comparable content differently. For a useful estimate, use the tokenizer or token-counting interface for the model you plan to call, then compare that estimate with the usage returned by the actual request: Gemini’s guide to understanding and counting tokens.

What can an AI agent run count?

An agent may make several model calls while handling a single task. Across those calls, usage can include input and output that never appears as a simple prompt-and-answer exchange. OpenAI’s agent documentation describes inputs such as instructions, tool definitions, conversation history, user messages, files or images, and tool results. Output can include the answer, tool-call arguments and reasoning. Some reasoning tokens are not shown in the final response, but OpenAI counts them as output usage and bills them at the output rate. OpenAI’s agent observability and usage guide also warns that cached input is still billed and repeated calls can process a large history.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tool calls and tool charges are different

A tool call can affect the bill in two ways. First, the model may generate arguments for that call, which count as model output. Second, using the tool may carry its own fee or billing unit. A tool’s returned content may also become input to a later model call. The treatment varies by provider and tool; Google’s Gemini pricing documentation distinguishes model inference from applicable tool charges and describes retrieved-content treatment for certain tools. Do not assume every tool call has one fixed token fee or that every tool handles retrieved content the same way.

Repeated calls, retries and handoffs add up

Agent loops can call a model again to interpret tool results, continue a task, retry work or hand it to another agent. OpenAI’s Agents SDK aggregates usage across model calls in a run, including calls that produce tool calls or handoffs; its usage records can include request count, input, output and total tokens, along with cached, cache-write and reasoning tokens. OpenAI notes that subagent calls can add usage. Google describes agent inference as standard model charges across inference loops, with applicable tools potentially charged separately. See the OpenAI Agents SDK usage guide and Google’s agents overview.

How token billing is calculated

For a token-priced service, the basic estimate is to multiply each billed token category by its applicable rate, then add the results. If rates are quoted per million tokens, divide each token count by one million before multiplying:

Estimated model charge = (input tokens ÷ 1,000,000 × input rate) + (cached input tokens ÷ 1,000,000 × cached-input rate) + (output tokens ÷ 1,000,000 × output rate)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is a calculation template, not a universal price formula. A provider may define usage categories and billing units differently. Add separately priced tools, storage, compute or other services where applicable. Rates can depend on the provider, model, endpoint, service tier, region and contract. OpenAI’s pricing page lists model- and category-specific terms; Google’s Gemini pricing page describes model inference and applicable tool pricing. Check the current page for the exact service you use rather than applying one provider’s rates to another.

Cached input is not automatically free: it remains input usage, though its rate may differ from uncached input. Reasoning tokens may be absent from the user-visible answer but still billed as output. Check whether your selected service also charges for cache creation or storage. When reviewing usage, make sure a cached-token figure is not added twice if it is already included in the reported input total.

Response length alone cannot predict a run’s total cost. Identical text can tokenize differently across models; repeated context and agent loops can increase usage; and input, cached input, output, reasoning and tool activity may have different rates. OpenAI advises comparing representative tasks, since a lower price per token does not necessarily mean a lower cost to complete the task.

How to check an agent’s usage and estimate its bill

  1. Identify the billing context. Record the provider, exact model, endpoint and service tier, plus whether the usage is covered by an API pay-as-you-go account, subscription, enterprise contract or another arrangement.
  2. Inspect the request or run’s usage records. Check response usage fields and, if applicable, the agent SDK’s per-call and per-run records. Names differ by endpoint: OpenAI documents prompt_tokens and completion_tokens for Chat Completions, and input_tokens and output_tokens for Responses. Some responses also expose cache and reasoning details. See OpenAI’s guide to reviewing API usage and costs.
  3. Total the whole workflow. Include every model call involved in the task, including calls that generate tool requests, retries, handoffs and subagent activity—not just the final answer’s call.
  4. Separate categories without double-counting. Where reported and priced separately, distinguish cached from uncached input and visible from reasoning output. Confirm whether cached usage is a subset of the input total before adding it in your calculation.
  5. Apply current rates and add other charges. Use the live pricing terms for the exact model and service, then include any applicable tool, storage or compute fees.
  6. Reconcile against billing. Compare your estimate with the provider’s usage dashboard or invoice. Usage fields may be best-effort, null or updated as accounting arrives; an interrupted stream may omit its final usage chunk. Missing usage data does not establish that a request used zero tokens.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to compare when choosing a model or plan

Test candidate options on the same representative task and compare both cost and the result achieved. Check the factors that can change the total:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Model, endpoint and service tier.
  • Input-to-output mix and how much context is repeated across calls.
  • Whether repeated context is eligible for caching, actual cache hits, and any cache-write or storage charge.
  • Reasoning or output usage and any relevant allowance or rate.
  • Number of agent loops, model calls, retries and subagents.
  • Tool charges and whether retrieved content is counted as model input.
  • Modality: text, images, audio, video or documents.
  • Billing arrangement, region, currency and current effective rate.
  • Quality, latency and successful task completion alongside total cost.

There is no stable cross-provider price comparison that applies to every agent run: rates and product terms change, and model- and tier-specific prices are not interchangeable. Use the provider’s current price page and your own representative usage records.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.