Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsReduce Claude API input costs by caching stable prompt content that is reused across requests, removing irrelevant or duplicated text, and checking actual token usage against the current price for your model. Caching can make repeated input cheaper, but it does not guarantee a hit or eliminate the initial cache-write charge. The savings depend on how often the same prefix is reused, when requests arrive, and how much of each prompt changes.
How prompt caching lowers repeated input charges
Anthropic’s prompt caching guide describes reusing previously processed prompt content across API calls. A request can reuse a recent cache entry when its prompt matches through the cache breakpoint. If there is no matching entry, Claude processes the content normally and may write the eligible prefix to cache.
Good candidates are content that is both substantial and genuinely repeated: stable system instructions, tool definitions, examples, reference documents, or recurring conversation context. Keep that reusable material together and put request-specific information after it. Changing content before the breakpoint can prevent a matching prefix.
Caching only affects eligible input tokens. You still pay for new or uncached input and for generated output. Minimum cacheable prompt lengths vary by model; content below the applicable minimum is processed without caching.
Recommended Free Tools
#1 Best Overall
Choose automatic caching or explicit breakpoints
Automatic caching
Automatic caching uses a top-level cache_control field and lets Anthropic manage the breakpoint as a conversation grows. It is a straightforward place to start when you want the system to handle breakpoint placement.
Explicit breakpoints
Explicit caching attaches cache_control to selected content blocks, giving you finer control over which sections are cached. Anthropic says one breakpoint at the end of stable content is sufficient for most cases. Multiple breakpoints—up to four—can help when sections change at different rates or a long conversation extends beyond the cache lookback.
Rank #2
In either approach, arrange content so stable instructions and reference material come before the breakpoint and per-request details come after it. Marking content for caching is not itself proof that a cache hit occurred; confirm it in message usage.
Pick a cache duration that fits the request pattern
Anthropic documents a five-minute default time-to-live (TTL) and a one-hour option. Its pricing documentation, accessed October 7, 2026, lists these general multipliers relative to the model’s base input price; model-specific exceptions apply.
| Cache event or duration | Documented price | Practical implication |
|---|---|---|
| Five-minute cache write | 1.25× base input price | Higher initial price for the eligible input written to cache. |
| One-hour cache write | 2× base input price | Higher write premium in exchange for the longer documented duration. |
| Cache read | Generally 0.1× base input price | Reads are cheaper than ordinary input; the live rate card lists exceptions, including Claude Fable 5.1 and Mythos 5.1 at 0.025×, and Opus 5.5 at 0.05×. |
Anthropic says the listed multipliers reach break-even after one cache read for the five-minute duration and two reads for the one-hour duration. Treat that as a rule of thumb for otherwise comparable cached tokens, not a guarantee for a workload: misses, changing prompt sections, model prices, and reuse timing all affect the outcome.
The TTL is measured from the start of the request that writes or reads the entry, and generation time counts against it. If a response takes several minutes, a five-minute entry may have little time left by the next request. Choose a duration based on when reuse actually happens, and compare the write premium with the reads you expect to get.
Rank #4
Rates vary by model and can change. The Claude Platform pricing page reports USD per million tokens with separate input, output, cache-write, and cache-read prices. As of October 7, 2026, it listed Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens; its five-minute cache writes were $2.50 per million and reads $0.20 per million. Check the live rate card for the model and cache option you use before estimating costs.
Shorten prompts without removing necessary instructions
Prompt trimming complements caching: caching makes repeated eligible input cheaper, while removing unnecessary text reduces the input sent in the first place. Review recurring prompts and remove material that is stale, duplicated, or irrelevant to the current task. Avoid repeating instructions or examples across turns when a stable reusable prefix can provide them once.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- Keep requirements explicit and relevant. Do not shorten instructions until the expected output becomes ambiguous; Anthropic advises using clear, specific instructions.
- Remove examples or documents that do not affect the current task, rather than blindly cutting every long section.
- Compare whether a shorter prompt still produces an adequate result. A prompt that triggers retries or extra follow-up requests can cost more overall even if its first input is smaller.
Use Anthropic’s token-counting endpoint to estimate input size for candidate requests before sending them. It accepts structured message inputs and can help compare prompt variants, budget costs and rate limits, or inform model routing. Anthropic cautions that actual message usage can differ slightly from the estimate. Token counting estimates prompt size; it does not test whether prompt caching will hit.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Verify savings with message usage, not estimates alone
Inspect the response usage fields to see what happened on real requests:
cache_creation_input_tokens: input tokens written to a cache entry.cache_read_input_tokens: input tokens retrieved from cache.input_tokens: input tokens after the last cache breakpoint that were not read from or written to cache.
Anthropic defines total input as the sum of those three values. When caching is enabled, input_tokens alone is not the entire prompt size. Token counting is useful before sending a request, but only actual message-response usage shows cache activity.
- Establish a baseline: record representative requests, model, input and output usage, and the applicable current prices.
- Apply one change at a time: test a reusable prefix or a shorter prompt on comparable tasks so you can identify which change affected usage.
- Check cache behavior: compare cache creation and read tokens across repeated requests, including whether the requests arrived within the selected TTL.
- Compare total cost and quality: account for ordinary input, cache writes, cache reads, output tokens, and any retries or follow-up calls. Keep the change only if it improves the real workflow without degrading results.
There is no universal percentage reduction: the result depends on the request pattern, amount of stable content, cache-hit rate, model, output, and current prices. Evaluate representative traffic rather than treating a lower estimated prompt size or a marked breakpoint as proof of savings.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

