Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no reliable universal price for an AI task in 2026. The cost depends on the model, input and output tokens, tools, repeated calls, and whether the work runs through an API or on your own infrastructure. To estimate it, measure a representative task, add every billed component, and divide the total spend by the number of tasks that meet your success criteria.

What “cost per task” should mean

A task should mean one successfully completed unit of work—not one API request. A single task may involve several model calls, search or code tools, retries, and intermediate agent steps. A price per million tokens is a metering rate; by itself, it does not say how much a particular task costs.

Keep three different figures separate:

  • API spend per task: the metered model and tool charges for all calls involved.
  • Cost per successful task: total spend divided by tasks that meet a defined quality or completion standard. This includes spend on failed attempts and retries.
  • Full system cost: API or infrastructure spend plus relevant operating and deployment costs, such as hosting, integration, monitoring, storage, and engineering.

How to calculate API cost for one task

For token-metered use, calculate the charges for every call made during the task:

Task API cost = input tokens × input rate + output tokens × output rate + applicable cache charges + tool and service charges

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use consistent units. If the provider lists a rate per million tokens, divide that rate by 1,000,000 before multiplying by token counts. Apply cache rates only to tokens the provider actually bills as cache reads or writes. Add separately billed search, code execution, or other services, then include every retry and intermediate agent call.

Illustrative calculation using a dated rate card

Google’s Gemini API pricing page lists Gemini 3.7 Flash Standard paid rates of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. At those rates, a hypothetical request using 10,000 input tokens and producing 2,000 output tokens costs $0.0075 for input plus $0.0075 for output, or $0.015 total before any applicable cache, grounding, or other service charges. This is an arithmetic example, not a typical task cost or a claim about model quality.

Gemini 3.7 Flash Standard paid rate listed by Google Input, per million tokens Output, per million tokens Effective period
Standard paid $0.75 $3.75 Through December 31, 2026
Standard paid $1.50 $7.50 Starting January 1, 2027

These are model-specific rates shown on Google’s pricing page, not general Gemini prices. Google also lists lower Batch and Flex rates and a higher Priority option for this model. Confirm the applicable rate, service tier, region, and date in the live pricing table before using an estimate.

Why provider rate cards do not give a universal answer

Rates differ by model and service tier, and providers may bill input, output, cached input, and tools differently. The task itself determines how many tokens and calls are needed. A fair comparison therefore holds the work and success standard constant rather than comparing unrelated headline rates.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI

OpenAI’s pricing page separates input, cached input, and output rates, with rates presented per million tokens. It also notes a 10% uplift for eligible regional-processing endpoints for models released on or after March 5, 2026. Select a specific model and applicable tier from the current pricing table; there is no single rate that describes all OpenAI tasks.

Google Gemini API

Google lists model-specific Standard and other service options, cache rates, and grounding charges. For Gemini 3.7 Flash, its pricing page says agent inference is billed at standard rates, including intermediate input, output, and reasoning tokens generated in agent loops. Grounding with Google Search can add a per-request charge after the listed free allowance.

Anthropic

Anthropic’s pricing page lists model-specific rates, prompt-cache write and read charges, a 50% saving for batch processing, and separate charges for tools including web search and code execution. Anthropic says its web-search charge does not include the input and output tokens needed to process requests, so include both parts when estimating a search-enabled task.

Anthropic’s model page estimates that Opus 5.5 costs 40% less to run than Opus 5 for typical token-billed workloads. That is Anthropic’s estimate, not an independent benchmark or a guarantee for a particular task. The page lists Opus 5.5 cache reads at $0.20 per million tokens.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How agents, tools, and retries change the bill

A task that appears to be one instruction can trigger multiple model calls and billed services. An agent may plan, call a tool, inspect its result, and then generate an answer. Each model step can add input and output tokens; search or code execution may be charged separately. A failed attempt can still incur charges even though it does not count as a successful completion.

Log these elements together for each task:

  • Input, output, and cached tokens for every model call.
  • Model, service tier, and relevant region for each call.
  • Tool calls and their separate charges, such as search or code execution.
  • Retries, failed runs, and intermediate agent steps.
  • Whether the final result met the task’s stated success condition.

A practical way to measure cost per successful task

  1. Define the unit of work. Specify what counts as one task and what result makes it successful—for example, a response accepted under your own quality criteria.
  2. Run representative examples. Use realistic task inputs and log all calls, token classes, tools, retries, and failures.
  3. Apply the matching rate card. Use the actual model, tier, cache behavior, and applicable rates for the period measured. Preserve any separate tool or service charges.
  4. Add the task’s full observed spend. Include failed attempts and retries in the spend rather than dropping them from the calculation.
  5. Divide by successful completions. Use total spend for the measured set divided by the number of tasks that met the success condition. Report the sample and conditions so the result is not mistaken for a universal price.

This gives an observed cost for a defined workload. If you need a full deployment estimate, add relevant hosting, integration, monitoring, storage, engineering, and operational costs rather than treating the API invoice as the cost of the whole system.

What self-hosting changes

Self-hosting replaces an API rate card with infrastructure economics; it does not make inference free. A GPU-hour quote alone omits whether the hardware is well utilized and may not capture the full capital and operating costs of serving valid work.

The 2025 paper Introducing LCOAI: A Standardized Economic Metric for Evaluating AI Deployment Costs proposes measuring total capital and operating expenditure per valid inference volume. LCOAI is a proposed metric, not an adopted universal standard, and the paper does not establish a universal cost per task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The 2026 paper Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation reports modeled effective costs of $0.21 to $15.25 per million output tokens on identical H100 hardware across its tested low-to-moderate enterprise loads of 1–10 requests per second. It reports underutilization penalties of 2.5–24× under those stated conditions and up to 36.3× near idle. These scenario-bound results illustrate how strongly utilization can affect unit cost; they are not universal market rates and should not be directly compared with API list prices.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare providers or deployment options fairly

Compare the same task under the same success criteria, rather than choosing a winner from headline token rates. Record the conditions that can change both cost and whether the result counts as useful work:

  • Effective cost per successful task: include all calls, retries, and failed runs.
  • Workload shape: input and output token volumes, context size, and cache reads or writes.
  • Tools and agent behavior: search, code execution, other billed services, and intermediate calls.
  • Model and task quality: decide what counts as success for your use case; a cheaper output that fails that standard is not an equivalent result.
  • Service conditions: note model, tier, region, latency, and availability requirements.
  • Deployment economics: for self-hosting, include utilization and amortized capital and operating costs per valid completion.

The current source set does not establish an apples-to-apples worked task comparison across OpenAI, Google, and Anthropic with identical capability, token mix, tools, success rate, and geography. A “cheapest provider” ranking without those controls would not answer what your task costs.

When an online cost calculator helps

Economize’s online LLM API Cost Calculator says it compares 197 models from 10 providers and asks for monthly input and output token volumes. Its page reports an update date of October 2, 2026. It can provide a first-pass estimate, but it cannot establish the token volume or number of calls your task will actually use. Validate calculator rates against provider pricing pages and replace assumed usage with measured traces.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep estimates current

Record the pricing date, model, service tier, region, token counts, cache behavior, tool charges, and success rate alongside each estimate. This matters because provider rates can change: for example, Google lists the Gemini 3.7 Flash Standard rate increase shown above as beginning January 1, 2027. Recalculate when the workload, service configuration, or applicable rate changes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.