The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To use prompt caching with Claude in Node.js, mark the stable part of a Messages API prompt with cache_control: { type: "ephemeral" }, then send later requests with the same prefix. Claude can reuse eligible processed input, which may reduce repeated input charges and improve time to first token for long prompts. The first request pays a cache-write premium, and caching does not make new input or generated output free.
What prompt caching does
Prompt caching lets Claude reuse eligible prompt content from an earlier API call when a later request has the same prefix through a cache breakpoint. It is most useful when an application repeatedly sends substantial shared context, such as system instructions, tool definitions, examples, long documents, or conversation history.
Anthropic describes the feature as reusing previously processed prompt portions to reduce costs and latency. The benefit is conditional: the shared prefix must meet the model’s minimum length, remain unchanged through the breakpoint, and still be within its cache lifetime. A cache read changes the billing and processing of repeated input; it does not remove the cost of request-specific input or generated output.
Choose automatic caching or a breakpoint
Automatic caching
Anthropic recommends starting with automatic caching for most use cases. Add a top-level cache_control: { type: "ephemeral" } to the request and the system manages a breakpoint as the conversation grows. Automatic caching follows the same eligibility thresholds, ordering rules, and lookback behavior as explicit breakpoints.
#1 Best Overall
Explicit breakpoints
Use a block-level breakpoint when you need to decide exactly which reusable section is cached, particularly when different prompt sections change at different rates. Place the breakpoint on the last stable content block, before request-specific material. Anthropic permits up to four breakpoints per request.
For example, stable system instructions can precede a changing user question:
Rank #2
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
apiKey: process.env.ANTHROPIC_API_KEY,
});
const response = await client.messages.create({
model: "CURRENT_CLAUDE_MODEL_ID",
max_tokens: 1024,
system: [
{
type: "text",
text: "Your stable, reusable system instructions go here.",
cache_control: { type: "ephemeral" },
},
],
messages: [
{ role: "user", content: "A request-specific question goes here." },
],
});
console.log(response.usage);
This illustrates the SDK call and cache-control structure; replace the model placeholder with a current model ID supported by your account. The example is not a tested program. The official TypeScript SDK package is @anthropic-ai/sdk, and its repository lists Node.js 20 LTS or later among supported runtimes. See the Anthropic TypeScript SDK repository and Claude prompt-caching documentation.
Put the breakpoint after the reusable prefix
A cache entry represents the prompt prefix through its breakpoint. Content that changes before the breakpoint changes that prefix and can prevent a cache hit. Put stable instructions, tools, examples, and shared documents first; place changing user-specific content after the breakpoint.
Rank #3
- Keep the reusable material identical across calls, including ordering and content.
- Keep relevant request settings consistent. Anthropic identifies changes to tool choice, image presence, thinking configuration, or output effort as possible cache invalidators.
- Ensure the prompt meets the minimum cacheable length for the selected model. Anthropic’s documented thresholds vary from 512 to 4,096 tokens across active models; check the current threshold for the model you use.
For the legacy Amazon Bedrock integration with Opus 4.6 and earlier, top-level automatic caching is not supported; use explicit block-level breakpoints instead.
Choose a cache lifetime based on request timing
Anthropic documents a default five-minute TTL and an optional one-hour TTL. The longer lifetime is useful when a repeat call may arrive after five minutes but within an hour. Both TTL choices behave the same for latency; the choice is about how long a cache can be reused and the write premium.
Rank #4
| Cache option | Write charge relative to base input price | Read charge relative to base input price | Anthropic break-even guidance |
|---|---|---|---|
| Five-minute TTL (default) | 1.25× | 0.1× standard multiplier | Pays off after one cache read |
| One-hour TTL | 2× | 0.1× standard multiplier | Pays off after two cache reads |
These are Anthropic’s standard documented multipliers, not universal dollar rates: some models have different multipliers, so check the live Anthropic pricing page for the target model. The break-even guidance compares cache write/read input costs with repeatedly paying base input price. Actual dollar savings depend on model price, prompt size, repeat frequency, TTL, and other pricing modifiers. Anthropic also says a five-minute cache refreshed while it is active continues to use it without another write premium, and cache hits are not deducted against the rate limit.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Verify cache writes and reads
Inspect the response’s usage object. cache_creation_input_tokens indicates input tokens written to cache; cache_read_input_tokens indicates cached input tokens read on a later request. Make repeated calls with the same model setup and stable prefix, then check that the usage fields reflect the expected write and read behavior.
Do not infer a particular latency gain or savings percentage from a cache hit alone. Anthropic says long documents generally see improved time to first token, but the actual outcome depends on the workload. Measure your own request timing and billing against an uncached baseline if you need a quantified result.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

