Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A KV cache stores the attention keys and values for tokens a decoder-only language model has already processed. Reusing that state avoids repeating work during token-by-token generation, but the cache takes memory and grows with active sequences. Once a model’s weights fit on the GPU, long contexts and many simultaneous requests can make KV-cache capacity—or the bandwidth needed to read the cache—a major limit on serving throughput. It is not a universal rule: the bottleneck depends on the model, workload, hardware, and serving setup.

What a KV cache stores

At each attention layer, a language model computes representations called keys and values for the tokens it processes. During generation, the model uses the current token’s query to attend to keys and values from earlier tokens. The KV cache retains those earlier key and value tensors so the model can reuse them at the next decoding step.

The cache is runtime state, not a copy of the prompt, the model’s weights, or the model itself. Weights are learned parameters loaded for inference; cached keys and values are derived from the tokens in active sequences. Hugging Face’s Transformers v5.3.0 Caching documentation describes the key/value cache and its growth as sequence length increases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why caching makes generation more efficient

Autoregressive generation produces output one token at a time. Without a cache, each new step would need to recompute attention state for the preceding tokens as well as the new one. With a cache, the model can reuse the earlier tokens’ keys and values and add state for the newly processed token. That reduces repeated computation, especially as the sequence grows, but requires storing and accessing the accumulated state.

Hugging Face’s Transformers v4.50.0 Optimizing inference documentation puts the repeated-work problem this way: “LLMs compute (key, value) (kv) values for each input token, and it performs the same kv computation each time because the generated output becomes part of the input.” The cache is the mechanism that lets generation reuse those values instead of recomputing them at every step.

How KV-cache pressure differs from weight pressure

Weights and the KV cache compete for memory, but they represent different limits. Weights are the model’s relatively fixed parameter footprint during inference. The KV cache is temporary state that grows as tokens are processed and as active sequences accumulate. A model may fit on a GPU but leave too little memory for long-context requests or a large number of concurrent requests. A very large model can instead be weight-limited before its cache becomes the main capacity concern.

Cache capacity and cache traffic are related but distinct. Capacity determines how much active state can fit. During decoding, the model also has to read cached state; moving that data can constrain speed even when it fits. Whether weights, cache capacity, cache reads, or another part of inference dominates depends on model dimensions and architecture, context and output lengths, concurrency, batching, cache data type, attention implementation, GPU memory bandwidth, and latency goals. The available documentation does not establish a universal point at which KV-cache costs overtake weight costs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the cache can limit throughput

Throughput is the amount of inference work a serving system completes over time. When a server handles requests concurrently, each active sequence needs its own applicable attention state. Longer contexts and more live requests therefore increase aggregate cache demand. If that state consumes too much of the available memory, the server may have to admit fewer sequences at once, limiting useful concurrency. If cache reads put pressure on memory bandwidth, decoding may also slow down.

This is why the claim that “the cache, not the weights” limits throughput is best understood as a common serving regime, not a law. It is plausible after weights fit and the workload has enough long or concurrent sequences to make runtime state substantial. It does not apply automatically to every model or request pattern, and cache size alone does not determine throughput.

Prefill and decode also have different characteristics. Prefill processes the input prompt; decode generates the output incrementally and reuses cached state. A change that helps one phase or one workload may not improve end-to-end performance for another. Compare systems using the workload and latency target that matter, rather than treating one cache setting or optimization as a universal speedup.

Ways serving systems manage KV-cache memory

Cache optimizations address different problems: how much GPU memory the cache occupies, how efficiently it is allocated, whether existing state can be reused, and how much memory a serving system reserves. The right choice depends on context length, output length, concurrency, and how often requests share prefixes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach What it changes Benefit Tradeoff or limitation
Keep the cache on the accelerator Retains active KV state in GPU memory. Avoids the extra movement associated with offloading. Uses memory that could otherwise support more cache state or other workloads.
Offload cache state Moves some cache state away from the GPU. Can free GPU memory for models or requests. May reduce generation throughput; Hugging Face’s cache-strategy documentation says the effect depends on the model and generation choices.
Paged cache allocation Organizes KV state in blocks rather than requiring one contiguous allocation per sequence. Can reduce allocation waste and support flexible sharing. Does not remove the memory needed for the underlying state or guarantee a fixed performance gain.
Automatic prefix caching Reuses KV blocks from earlier requests when prompt prefixes match. Avoids redundant work for shared prefixes. Helps only when requests have reusable matching prefixes; vLLM documents this feature as Automatic Prefix Caching.
Increase the cache-memory budget Allows the serving engine to use more memory for cache capacity. May let the system support more cache state and improve throughput capacity. An excessive allocation can cause out-of-memory errors; vLLM’s LLM API documentation describes this capacity-versus-OOM tradeoff.

These are not interchangeable controls. Offloading trades memory placement against data movement; paging addresses allocation efficiency; prefix caching exploits repeated input; and a larger budget changes how much memory is made available. NVIDIA’s TensorRT-LLM KV Cache System documentation also describes cache reuse, offloading, eviction, and allocation controls. Exact features and option names vary by serving engine and version, so consult the documentation for the version actually deployed.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the PagedAttention result does—and does not—show

The authors of the 2023 PagedAttention paper reported 2–4× higher throughput at the same latency level on their evaluated workloads compared with the systems they tested, including FasterTransformer and Orca. The result supports the value of the paper’s memory-management approach under those experimental conditions; it is not a guaranteed gain for every current vLLM deployment, model, or workload. It should not be read as a general multiplier for KV caching itself.

How to reason about a workload

When cache memory appears to constrain a serving system, first distinguish a capacity problem from a speed problem. A capacity problem means active state limits how many or how long the requests can be. A speed problem may involve the time needed to move cached state, but the sources cited here do not define a universal bandwidth threshold. The distinction matters because freeing memory, reducing allocation waste, and reducing repeated computation address different parts of the system.

  • If long inputs or many concurrent sequences are central: account for the cache’s growth with active sequence length and concurrency when evaluating how many requests can fit.
  • If prompts often share prefixes: prefix reuse may avoid redundant work for matching portions, as described in vLLM’s Automatic Prefix Caching documentation.
  • If GPU memory is the immediate constraint: compare keeping cache state on the GPU with the available offloading options, weighing freed capacity against possible throughput degradation.
  • If increasing the memory budget: treat it as a capacity adjustment, not a guaranteed speed setting; an excessive reservation can produce an out-of-memory error.
  • If comparing reported throughput: keep the model, workload, latency target, concurrency, cache precision, and engine implementation in view. A result from a different evaluated setup may not predict yours.

There is no single cache-size figure or universal crossover between weight and cache costs established here. Model architecture, KV precision, sequence lengths, concurrency, engine behavior, and hardware all affect the result, so model-specific sizing requires those details.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.