Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI API bills usually separate the tokens you send from the tokens a model generates, and may price eligible cached input differently from uncached input. To estimate a request, count each billable category separately, apply the rate for the exact model and usage mode, then add any applicable cache-storage or feature charges. Rates and caching rules differ by provider, model, tier, and modality.

What do input and output tokens mean on an API bill?

Input tokens are the content supplied in a request; output tokens are content generated by the model. Providers can charge different rates for each category. OpenAI’s pricing page lists separate input, cached-input, cache-write, and output rates, with additional distinctions for some models and context lengths. Google likewise publishes rates by model and usage configuration on its Gemini API pricing page.

Output usage may include generated reasoning tokens that are not shown in the final response. OpenAI’s token guidance explains that reasoning tokens can count as output and be billed even when they are not visible to the user. A short-looking answer therefore does not necessarily mean that little output was generated.

How do you calculate the token charge?

Use this planning equation, then confirm the provider’s definitions and the usage reported for your request:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Estimated token charge = uncached input tokens × uncached-input rate + cached input tokens × cached-input rate + output tokens × output rate + applicable cache-write, cache-storage, or feature charges.

If rates are quoted per million tokens, divide each token count by 1,000,000 before multiplying by its rate. Do not automatically add a cache-write fee: OpenAI documents cache-write pricing as an alternative input-token rate, not a fee that is always added on top of uncached input.

Worked arithmetic example

For a hypothetical request with 10,000 input tokens and 1,000 output tokens, calculate the input charge using the selected model’s input rate and the output charge using its output rate. Add the two. For a real estimate, first split eligible cached input from uncached input and include any applicable storage or non-token charges. This shows the calculation method; it is not a quote of actual spend.

Why token counts are not word counts

Tokens are units a model uses to process text, not a fixed number of words. OpenAI’s Help Center offers rough English-language estimates: about four characters per token, about three-quarters of a word per token, or roughly 75 words per 100 tokens. These are approximations, not universal conversion constants. Results vary with language, spelling, capitalization, spaces, and the model’s tokenizer. A plain-text estimate can also miss message structure, tool definitions, schemas, images, or files.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For better estimates, use the tokenizer appropriate to the model and inspect API-reported usage. OpenAI explains token counting in its token guide.

When can caching lower API costs?

Caching can help when a substantial part of your prompt or corpus repeats across requests. Instead of processing every repeated token at the ordinary input rate, a provider may charge a lower rate for eligible cached input. Whether a request qualifies—and whether cache writes or storage cost extra—depends on that provider’s implementation.

OpenAI prompt caching

OpenAI says cache reuse depends on the rendered prompt prefix matching, while eligibility and cache breakpoints depend on the model. Its current guide specifies a 1,024-token minimum cacheable prompt for GPT-5.6 and later, with different thresholds for earlier models. Check the current prompt-caching guide for the model you use.

As a model-specific illustration in that guide, GPT-5.6-and-later cache writes are priced at 1.25 times the standard uncached input rate; cache reads are priced at 0.1 times that rate for most of those models and 0.05 times for GPT-6.1 Sol. At a 0.1-times read rate, one write plus nine full reads costs 2.15 times the ordinary input cost of one processing pass, versus 10 times that cost for ten uncached passes. This illustrates the listed rates for those models; it does not establish savings for other prompts or providers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Gemini context caching

Google describes implicit caching for Gemini 2.5 and newer models, as well as explicit caching as a separate feature. For explicit caching, cost depends on token count and time-to-live (TTL); the documented default TTL is one hour when unset, and storage duration can contribute to the bill. Cached-token, uncached-input, and output charges may all apply. Google’s caching documentation identifies the feature as Beta and says the relevant endpoints and SDK methods are under v1beta; check its context-caching documentation for current status and implementation details.

Google notes that “At certain volumes, using cached tokens is lower cost than passing in the same corpus of tokens repeatedly.” The qualification matters: savings depend on the volume and the caching costs, including storage.

What do published prices look like?

Rates are snapshots, not permanent market prices. OpenAI’s pricing page is organized by model and quotes rates per million tokens; columns can distinguish input, cached input, cache writes, output, and short or long context. Check the row for the model and configuration you intend to use rather than treating any one rate as an OpenAI-wide price.

One Google example illustrates why modality and caching need separate rows. On the Google AI for Developers pricing page accessed October 7, 2026, Gemini 3.1 Flash-Lite Standard was listed at $0.25 per million text, image, or video input tokens; $0.50 per million audio input tokens; $1.50 per million output tokens; and $0.025 per million text, image, or video cached tokens, plus $1.00 per million tokens per hour for storage. These are configuration-specific figures from that date, not a general Gemini rate. The page also lists different rates for Batch, Flex, and Priority tiers. Consult the current pricing page before estimating a bill.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some entries on Google’s page have rates that change on a stated future date. Keep the model, tier, modality, and effective date together when comparing them; do not combine a current rate with a later rate.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should you compare two APIs for your workload?

Compare the cost of completing the same representative task, not just the headline price per million tokens. OpenAI warns that models may tokenize identical text differently and generate different amounts of output or reasoning; it advises testing representative tasks rather than comparing only visible response length.

  1. Match the task and model capability. Compare models that can perform the work you need, not provider averages or unrelated model tiers.
  2. Describe the request mix. Estimate prompt size, expected completion size, and any reasoning usage the API reports.
  3. Check the exact rate row. Match model, service tier, modality, context length, and effective date. Account for different units or charges for audio, images, video, batch processing, priority service, grounding, or other features where applicable.
  4. Understand caching before counting savings. Find out whether caching is implicit or explicit, what content must match, minimum eligible size, cache lifetime, read and write rates, and any storage charge.
  5. Run representative requests and inspect usage. Use API-reported counts for each category, calculate the cost per completed task, then multiply by expected volume. A token-rate comparison alone cannot establish which option produces the lower total bill.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.