Recommended Free Tools
There is no universally cheapest provider for cached prompts. Compare the full cost of your workload—not just the cached-token rate—including cache creation, cache reads, storage time where billed, uncached input, generated output, and the cache hits you actually get. Provider rates and model availability change, so use the current rate cards for the exact models and service tiers you plan to run.
What to compare before choosing a provider
Cached-prompt pricing is not a like-for-like line item across Anthropic, OpenAI, and Google Gemini. Their published schedules describe cache costs differently, and total cost depends on how often a prompt prefix is reused and how much output each request generates. Set the workload first, then compare the relevant rates.
- Model and service tier: Select the exact model, context class, and tier. Pricing may vary across models or tiers.
- Prompt pattern: Record the reusable prefix size, how often requests repeat it, and the timing between repeats.
- Cache costs: Include creation or write charges, reuse charges, and any storage-duration charges.
- Other tokens: Count ordinary input that is not served from cache and the output tokens generated.
- Real cache performance: Use observed cached-token usage where available, rather than assuming every repeated request will hit the cache.
Use each provider’s current official rate card: Anthropic pricing, OpenAI API pricing, and Gemini Developer API pricing.
How each provider prices cached prompts
Anthropic Claude API
Anthropic’s pricing documentation separates base input, cache writes by duration, cache reads or refreshes, and output. Its published rate-card multipliers state that 5-minute cache-write tokens cost 1.25 times the base input rate, 1-hour cache-write tokens cost 2 times that rate, and cache-read tokens cost 0.1 times the base input rate. These are pricing multipliers, not a complete per-request cost: the chosen model’s base rate and the workload’s mix of writes, reads, and uncached tokens still matter.
#1 Best Overall
Because write rates differ by cache duration, include the intended cache lifetime and reuse pattern in any estimate. The model prices shown on a provider page can change; check the current table for the model you intend to use rather than applying these multipliers to an older model’s listed rate. Source: Anthropic pricing documentation.
OpenAI API
OpenAI publishes model-specific rates for input, cached input, cache writes, and output; rates can also differ by context class. Its prompt-caching guide describes reuse of a matching prompt prefix, but keeping a session open does not guarantee a cache hit. Check usage information and base estimates on cached-token activity observed for your request pattern.
Do not equate OpenAI’s cached-input price with another provider’s cache-read price without checking whether cache creation and any separate storage charges are included on the same basis. Consult the current API pricing schedule and prompt-caching guide.
Google Gemini API
Google’s Gemini pricing documentation lists context-caching token rates and, for paid tiers in the schedule reviewed, separate storage charges per million tokens per hour. The schedule includes a $0.50 per 1,000,000 tokens per hour storage entry for some paid-tier listings, but other entries have different rates or tier terms. Treat that figure as a specific schedule example, not a universal Gemini price; confirm the current model, tier, and applicable terms before using it in a budget.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Google documents implicit-cache usage reporting, while its explicit-caching guide explains how cached content can be reused in later requests. Verify that the model supports the caching mode and any threshold your workload requires. See Gemini pricing, context caching, and explicit caching for the Generate Content API.
Build a workload-level cost estimate
For each provider, estimate cost over the same time window and request volume. Use the same prompt, expected output, model class, and service tier where practical. A useful accounting structure is:
Rank #4
- Ordinary input tokens multiplied by the applicable standard input rate.
- Cache creation or write tokens multiplied by the applicable write rate.
- Cached input tokens multiplied by the applicable cache-read or cached-input rate.
- Cache storage multiplied by the applicable duration charge, if the provider bills it separately.
- Output tokens multiplied by the applicable output rate.
Estimate at least two scenarios if cache-hit performance is not yet known: a conservative case with fewer cache hits and a more optimistic case based on plausible reuse. Replace the assumptions with measured usage once the workload is running. OpenAI specifically recommends checking usage information because a matching prefix does not guarantee a hit; Google documents cache-usage reporting. See the OpenAI prompt-caching guide and Google context-caching guide.
There is no universal break-even threshold established by the providers’ pricing pages. The result depends on current model rates, write and storage costs, how long content remains reusable, actual hit rates, and output volume.
Best Value
Why a cached-token price alone can mislead
- Cache writes can change the economics: A low read rate does not by itself show whether creating the cache pays off for your reuse pattern.
- Storage may be billed separately: Gemini’s published schedule can add an hourly storage charge on paid tiers.
- Not all repeated prompts hit: OpenAI documents prefix matching but does not guarantee a hit simply because a session continues.
- Uncached input and output remain in the bill: A workload with substantial new input or generated output can cost more than the cache rate suggests.
- Provider labels are not identical: Compare what each line item includes before treating two cached-token rates as equivalent.
When the comparison is ready to make
A defensible provider choice needs a defined workload and current prices for its exact model and tier. Record the ordinary and cached input, write or creation events, cache duration, storage where charged, output, and observed hit rate. Then compare total estimated cost rather than ranking providers by a single cached-input number.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

