Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no universal price for an AI request. For a token-priced hosted API, calculate the charge from the request’s billed input, output, and any separately priced usage categories. For a self-hosted model, divide the serving costs you choose to include by the number of completed requests served over the same period. In either case, use representative traffic and compare options against the same workload and service requirements.

Calculate the charge for a token-priced API request

For each billed usage category, multiply its token count by that category’s price per million tokens, then divide by 1,000,000. Add the resulting charges:

Request model charge = Σ(category tokens ÷ 1,000,000 × category price per million tokens)

At minimum, keep input and output separate because they may have different rates. Add separate terms for cached input, cache writes, or other features when the provider’s rate card bills them separately. OpenAI’s published enterprise formula, for example, adds input, cached-input, and output charges; it does not set one universal fee per request. See OpenAI API pricing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Arithmetic example

For a request with 2,000 input tokens and 500 output tokens, at rates of I dollars per million input tokens and O dollars per million output tokens, the charge is 0.002 × I + 0.0005 × O. This illustrates the calculation; it is not a quoted price or measured benchmark. If some input qualifies for a cache-read rate, split those tokens out and apply that rate rather than charging all input as ordinary input. Anthropic documents separate cache-write and cache-read categories in its API pricing.

Estimate a service’s average cost, not just one request

A single hand-picked prompt is rarely representative of a service whose context length, response length, cache behavior, or tool use varies. Group requests into meaningful classes, calculate each class using its actual billed usage, and weight the results by each class’s share of traffic.

  1. Identify the billing route. Record the provider, model, endpoint, region, and service tier your workload will use. Direct API rates may differ from cloud-platform rates or regional rules. AWS says OpenAI models on Bedrock are billed through AWS; Anthropic says pricing for partner-operated Bedrock and Vertex AI is independent of its direct API regional pricing. Check the applicable OpenAI price list, Anthropic price list, or Amazon Bedrock pricing.
  2. Measure representative usage. Use provider usage fields or request logs to capture input and output tokens for typical and unusually long requests. Character counts are not a reliable substitute when token counts are available.
  3. Separate priced categories. Track cache writes, cache reads, and other billable features separately wherever the rate card assigns them distinct prices. Cache eligibility and duration are provider-specific.
  4. Include relevant modifiers. Check whether batch processing, priority or fast service, long-context brackets, geographic processing, or tool charges apply. Do not assume discounts or modifiers combine; verify the provider’s terms for the route you will use.
  5. Weight by the observed request mix. Calculate costs for classes such as short and long prompts, typical and long completions, cache hits and misses, and tool-using requests. Multiply each class’s cost by its share of requests, then add the weighted costs.
  6. Scale to expected volume. Multiply the weighted average cost per request by the expected number of requests in the period. Keep a range if traffic volume, output length, or cache behavior is uncertain, and recheck the official rate card before relying on the estimate because rates and service tiers can change.

Estimate self-hosted cost from completed work

For a self-hosted model, choose the costs your estimate is intended to include, allocate those serving costs over a defined period, and divide by the completed requests served in that period:

Self-hosted cost per completed request = allocated serving cost for a period ÷ completed requests served in that period

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The allocation may include rented or amortized accelerators and associated operating costs, depending on your accounting boundary. Measure throughput with the intended model, real workload, concurrency, and latency target. Include paid capacity that sits idle: infrastructure expense persists when output falls, so low utilization can raise effective cost per token or completed request. NVIDIA’s TCO guidance notes that hourly hardware price alone obscures throughput and latency.

Compare deployment options on equal terms

Do not compare an API bill directly with a GPU hourly quote. Estimate cost per completed request or cost per token for the same workload and service requirements, then assess the trade-offs that matter to your service:

  • Model capability and quality on the task.
  • Input, output, and cache-adjusted prices for the observed request mix.
  • Latency and throughput at the required concurrency.
  • Batch, regional, and service-tier modifiers.
  • For self-hosting, operational overhead and utilization.

NVIDIA’s sizing guidance identifies model selection, request lengths, cache hit rate, concurrency, latency targets, and contract duration as relevant sizing inputs. The right choice depends on your workload; the available pricing and sizing factors alone do not determine it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret published inference-cost benchmarks

NVIDIA’s AI inference page reports $4.20 per million tokens for its stated Hopper configuration and $0.12 per million tokens for its stated Blackwell configuration. These are vendor-presented benchmark claims tied to particular hardware and test conditions, not general market prices or a substitute for benchmarking your model and traffic pattern.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA’s TCO guidance states: “AI inference economics depend on the cost per token and overall system throughput rather than raw hourly hardware rates.” Treat this as NVIDIA’s framing, not an independent standard.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.