What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Agentic systems often resend the same long prompt prefix—system instructions, tool definitions, reference material and conversation history—on multiple model calls. When a provider recognizes that prefix as a cache hit, it can charge less for those cached input tokens and avoid repeating much of the work to process them. Across an agent loop, that difference can add up. But savings depend on the model, provider, matching prefix and time between calls; a cache hit is not automatic, and it does not discount new input or generated output.
What cache-hit pricing means for an agent
A prompt prefix is the beginning portion of a request that remains identical across calls. A cache hit means the provider can reuse its processed state for a matching prefix; a cache miss means that reuse is unavailable and the input is processed at the ordinary rate. In OpenAI’s description, prompt caching reuses the matching prefix’s key-value state, while new input still has to be processed (OpenAI prompt-caching guide).
In an agent loop, a model may receive the same instructions and tool schemas, choose an action, wait while a tool runs or a person approves it, then receive another request containing the prior context plus new results. If the reusable beginning of that request hits the cache, the cached portion can cost less than uncached input. The new tool result, changed user message and generated response still have their own processing and billing.
The effect compounds because agents can make many calls for one task. A modest per-call discount on a large repeated prefix may matter more than the same discount on a single short chat prompt. Conversely, a large advertised cache discount does little for a workload whose prompts rarely match or whose waits exceed the cache’s retention window.
#1 Best Overall
How the pricing math works
Compare the cost of writing a prefix to the cache with the cost of reading it on later calls. A write may cost more than a standard input pass; a read may cost much less. The write premium is recovered only if enough subsequent calls reuse the cached prefix before it expires or becomes ineligible.
| API pricing example | Cache write | Cache read | What the figures mean |
|---|---|---|---|
| OpenAI, GPT-5.6 and later | 1.25× standard uncached input rate | 0.1× on most such models; 0.05× on GPT-6.1 Sol | Model-specific multipliers in OpenAI’s guide; check the API pricing page for the model’s current prices. |
| Anthropic Claude API | 1.25× base input price for a 5-minute cache; 2× for a one-hour cache | Generally 0.1× base input price, with model-specific exceptions | Anthropic’s documented Claude API rates; partner platforms such as Bedrock and Google Cloud may price independently. |
For OpenAI’s illustrated 0.1× read rate, one write plus one full cached read costs 1.35 times the cost of one ordinary input pass. Two ordinary passes cost 2 times that pass. One write plus nine reads costs 2.15 times one pass, compared with 10 times the pass for ten uncached requests. These examples isolate the repeated prefix’s input-token cost: they do not include new tokens, output, or other platform charges. The guide’s multipliers are not universal across OpenAI models.
Rank #2
Anthropic says its general 0.1× read rate makes a 5-minute cache write’s premium pay back after one cache read, and a one-hour write’s premium after two reads. That comparison concerns token rates; other billed input, output and platform charges still apply. Model exceptions and platform-specific pricing can change the calculation (Anthropic Claude pricing).
For a workload estimate, treat the initial write and later reads separately: count the cached prefix tokens, how often a follow-up request reuses them, and how many of those follow-ups actually hit. Then add uncached input, new input, output and applicable platform charges. The headline read discount alone is not the total cost of an agent task.
Rank #3
Why agent loops are especially sensitive to cache misses
Tool calls create repeated context
An agent frequently needs its governing instructions and tool definitions on every model call. Those elements can form a valuable shared prefix when kept unchanged. OpenAI recommends preserving conversation history and keeping tool definitions stable; changing content earlier in the prompt can prevent later content from matching the same prefix (OpenAI prompt-caching guide).
Tool runs and approvals create gaps
An agent’s next call may not arrive immediately. A tool can take time to finish, or a human approval can pause the loop. If that gap exceeds the provider’s cache lifetime, the next request may be a cache miss even when the prompt text is otherwise unchanged. A 2026 preprint by Maxim Khailo analyzes this “think, act, wait” pattern and the economics of periodic keepalives. Its analysis is one researcher’s study, not an official provider recommendation or a universally validated operating rule (Khailo, “Keeping the Cache Warm Pays”).
Matching and routing matter
A cache hit requires more than sending similar-looking text. The provider’s matching rules, cache availability, routing and retention settings can affect reuse. OpenAI describes machine-local cache states and says cache location, routing, lifetime and traffic can influence whether a request hits. Its newer GPT-5.6-and-later guidance documents explicit cache breakpoints and a 30-minute retention control, with at least 30 minutes after the latest write or reuse for that generation. Older OpenAI models have different behavior, minimum lengths and retention options, so those newer rules should not be assumed for every model.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteWhat to keep stable to improve reuse
- Put reusable instructions, reference material and tool definitions at the beginning of the prompt, with per-turn content later where the API permits.
- Keep tool schemas and other early prompt content stable instead of rebuilding or reordering them on every call.
- Append conversation history rather than rewriting its earlier prefix when the application design allows it.
- Check each provider and model’s minimum cacheable length, breakpoint behavior and retention settings; do not transfer one provider’s rules to another.
- Account for the actual time between an agent call and its follow-up, including tool execution and approval waits.
These practices improve the chance that a later request can reuse a prefix; they do not guarantee a hit. OpenAI’s September 22, 2026 announcement describes GPT-6 prompt caching as designed for persistent agents and says eligible shared prefixes reused within a 30-minute window can receive discounts of up to 90% on cached input tokens. “Up to” is important: realized savings depend on the workload and model (OpenAI’s GPT-6 announcement).
Best Value
How to tell whether caching saves your workload money
- Choose a representative agent task. Include ordinary turns, tool calls and realistic approval or tool delays rather than estimating from a single short prompt.
- Identify the reusable prefix. Separate stable instructions and schemas from new user messages, tool results and other changing content.
- Calculate the write premium and likely reads. Use the exact provider, model, API platform and cache-retention mode. Estimate how many follow-up calls can reuse the prefix before it expires.
- Inspect actual usage and billing. Use the provider’s usage details or dashboard to measure cached input tokens, cache writes and reads on representative runs. Compare total input and output costs, not just the published cache-read rate.
- Recheck after changing the agent. Prompt edits, tool-schema changes, history rebuilding, routing or a different delay pattern can alter the match rate and the cost calculation.
OpenAI’s announcement attributes a company result to GitHub Chief Product Officer Mario Rodriguez: across billions of requests to OpenAI models, GitHub reduced by more than 50% the share of prompt tokens requiring fresh processing relative to its previous baseline. This is GitHub’s reported experience, not an independent study or a forecast for other agent workloads. Rodriguez said, “OpenAI’s prompt caching plays a critical role in helping GitHub Copilot deliver fast, efficient experiences at scale” (OpenAI, September 22, 2026).
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

