Free tools Windows power users keep installed
One-click scans. No signup required.
Anthropic prompt caching can lower API input costs when requests reuse the same prompt prefix. Put a cache breakpoint after stable content, keep changing material after it, and verify reuse in the response’s cache-usage fields. A cache marker alone does not guarantee savings: the prompt must meet the model’s minimum cacheable length, the prefix must match, and requests must arrive within the selected cache lifetime.
How prompt caching works
Prompt caching lets Claude reuse a matching prefix of a request across API calls instead of processing that repeated content as ordinary input each time. Suitable material includes system instructions, tool definitions, examples, documents or images in user turns, and earlier tool-use or tool-result content. It is most useful when a large body of context stays unchanged while the user’s latest message or other request details vary.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The Claude AI Advanced Handbook: Model and Effort Economics for Claude Opus 5: Real Cost Per Task,... | $9.99 | Buy on Amazon |
Anthropic describes the feature as reusing previously processed prompt portions across API calls to reduce cost and latency. See Anthropic’s prompt caching guide and its pricing documentation for current configuration and rates.
Choose automatic caching or explicit breakpoints
Automatic caching
For the simplest starting point, add cache_control: {"type": "ephemeral"} at the request’s top level. Anthropic says this automatically places the breakpoint on the last cacheable block and moves it as conversation history grows.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Explicit breakpoints
Place cache_control on selected content blocks when you need to control exactly which prefix is cached. This can help when, for example, system instructions remain stable but retrieved context changes between requests. Anthropic supports up to four breakpoints; adding them does not itself add a charge. Billing depends on the content written to and read from cache.
In either approach, put a breakpoint after the last block that stays identical from request to request. Place timestamps, incoming user messages, and other changing content after the stable prefix. If cached content or relevant request settings change, some or all of the prefix may no longer match.
Choose a cache lifetime and assess the cost
The default cache lifetime is five minutes. Anthropic measures it from the start of the request that writes or reads the entry, so time spent generating a long response counts against that window. Reuse refreshes the cache without an additional charge. A one-hour lifetime is also available for an additional write premium and can be useful when requests are spaced more than five minutes apart but still arrive within an hour.
| Decision factor | Five-minute TTL | One-hour TTL |
|---|---|---|
| Standard cache-write price | 1.25× the model’s base input price, according to Anthropic’s pricing checked October 7, 2026 | 2× the model’s base input price, according to Anthropic’s pricing checked October 7, 2026 |
| Standard cache-read price | 0.1× base input price, subject to model-specific exceptions | 0.1× base input price, subject to model-specific exceptions |
| When it may fit | Requests reuse the prefix within five minutes | Reuse gaps exceed five minutes but remain within an hour, or operational needs justify the higher write cost |
| Main trade-off | A long response leaves less time for the next request to reuse the entry | Higher write premium; the longer lifetime should be useful often enough to justify it |
These are Anthropic’s standard multipliers, not a promise of a particular saving. Rates are model-specific and can change; the pricing page may also list exceptions to the standard cache-read multiplier. Check the current rate for the model and platform you use before estimating costs.
At the standard 0.1× read multiplier, Anthropic says a five-minute cache write can break even after one cache read, while a one-hour write can break even after two reads. This compares the write premium with reads priced at 10% of base input cost; actual results depend on prompt size, model pricing, cache hits, TTL, and when entries expire.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Set up caching and verify reuse
- Identify the repeated prefix. Look for substantial content that recurs unchanged, such as system instructions, tool definitions, examples, or a long document.
- Choose a breakpoint strategy. Start with automatic caching for a straightforward request or conversation. Use explicit breakpoints if stable and changing sections need separate boundaries.
- Keep variable content after the breakpoint. Put request-specific items such as timestamps and the new user message after the stable prefix. Avoid changing cached content or relevant request settings if you expect a cache hit.
- Select the TTL based on request cadence. Use the five-minute default when reuse normally happens within that window. Consider one hour only when the longer window is useful enough to offset its higher write premium.
- Inspect the response usage fields. Check
cache_creation_input_tokensandcache_read_input_tokens. Anthropic defines total input asinput_tokens + cache_creation_input_tokens + cache_read_input_tokens;input_tokensalone represents only the uncached portion after the last breakpoint. - Investigate zero cache counts. If both cache-creation and cache-read counts are zero, check whether the prompt meets the model’s minimum cacheable length and whether a change invalidated the prefix. Minimum lengths vary by model, so confirm the current requirement in the documentation.
Account for timing and platform differences
A cache entry becomes available after the first response begins. Concurrent requests sent before that point may not receive a cache hit, so do not assume a group of simultaneous first requests will all reuse a newly written entry.
Anthropic’s documentation lists active Claude models and the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry as supported. Cache minimum lengths, usage field names, and setup instructions may differ by model or hosting platform. Follow the provider-specific instructions for the deployment you use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

