Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tokenization turns text into a sequence of vocabulary IDs; a Transformer turns those IDs into representations that attention can compare and combine. During autoregressive generation, a key-value (KV) cache keeps the attention keys and values already computed for earlier tokens, so the model can reuse them instead of rebuilding the entire prefix at every step. That saves repeated work, at the cost of memory that grows with the cached context.

What tokenization does

A language model does not normally receive a sentence as raw text. A tokenizer breaks it into units from a fixed vocabulary—often subwords produced by methods such as byte-pair encoding (BPE) or WordPiece—and maps those units to integer IDs. The model then looks up or constructs a vector representation for each ID.

Segmentation depends on the tokenizer’s vocabulary and rules. A familiar word might become one unit, while a rare word may be split into several pieces. Spaces, punctuation, scripts, and unusual character sequences can also affect the result. There is no universal number of tokens per word: a different tokenizer can produce a different sequence for the same text.

This matters because sequence length affects model work and memory. A longer token sequence means more positions to process, and during generation it can mean a larger attention cache. Tokenization is therefore a fundamental preprocessing step, not just a display or formatting choice. Song and coauthors’ 2020 Fast WordPiece paper evaluated tokenizer speed in a particular general-text setting; its reported speed comparisons should not be treated as a measure of current LLM serving performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How tokens become attention vectors

After token IDs are converted into vectors, Transformer layers update those representations using attention and other operations. Vaswani and coauthors introduced the Transformer in 2017 as an architecture based solely on attention mechanisms, dispensing with recurrence and convolutions. In attention, each position uses a query to compare against keys; the resulting weights determine how it combines the corresponding values.

Query, key, and value

At a high level, a layer derives three vectors from its input representation using learned projections:

  • Query (Q): what the current position is seeking from other positions.
  • Key (K): information used to judge how relevant a position is to a query.
  • Value (V): the content that gets mixed together according to the attention weights.

A query is compared with keys to produce scores. The model normalizes those scores into weights, then uses the weights to form a weighted combination of values. For causal language models, a position cannot use information from future positions; this restriction lets the model predict the next token without seeing it in advance.

These are not three separate meanings assigned to the original words. They are learned, layer-specific projections of the evolving token representations. As information passes through layers, each position can incorporate relevant information from earlier positions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What happens during prompt processing and generation

Prefill: process the prompt

When a prompt arrives, the model processes its token sequence in a prefill phase. It computes layer representations for the prompt positions, including the keys and values needed by attention. Causal masking ensures each position only uses permitted earlier context. Implementations can process prompt positions together rather than generating them one at a time.

Decode: generate one token at a time

In autoregressive decoding, the model predicts a next token, appends it to the sequence, and repeats. For each new position, a layer computes the new position’s query, key, and value. The query can attend to keys and values for the preceding context as well as the new position, subject to the model’s attention rules. The chosen next token becomes part of the context for the following step.

Why a KV cache makes decoding faster

Without a cache, an implementation may recompute attention keys and values for the whole prefix at every decode step. A KV cache stores the keys and values already derived by attention layers for earlier tokens. When another token is generated, the model computes and adds that token’s states, then reuses the stored states for the prefix. Hugging Face Transformers documentation describes this as eliminating repeated work by storing key-value pairs from previously processed tokens.

The cache does not eliminate all work: each new token still has to pass through the model, and its query must attend to the available context. It avoids recalculating the older positions’ cached key and value states. This is particularly useful as the prefix grows. The trade-off is that cache storage increases with the number of retained tokens; exact memory use and speed depend on the model, precision, context, batch, and implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to estimate KV-cache memory

A useful first-order estimate for an uncompressed cache is:

bytes ≈ 2 × layers × batch size × cached tokens × KV heads × head dimension × bytes per value

The factor of 2 accounts for storing both keys and values. “KV heads” means the number of key/value heads actually stored per layer, which may be smaller than the number of query heads in architectures that share key/value heads. For example, if values use a two-byte representation, the estimate counts two bytes for each stored key element and two for each stored value element.

This is a tensor-size estimate, not a guaranteed device-memory reading. Quantization may change the effective storage and add metadata; implementations can also have allocation overhead, padding, or other runtime memory use. Sliding-window layers may retain fewer than all prior tokens, whereas an unbounded-attention layer’s cache grows with the context. Batching multiplies the cache requirement when each batch item has its own sequence state. Since model dimensions and cache policies differ, there is no single universal KV-cache size.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Dynamic, static, quantized, or offloaded cache?

Cache choices trade memory use against decode performance, compilation support, compatibility, and operational complexity. The right choice depends on the model and workload rather than a universally best setting.

Cache type How it works Benefits Costs and cautions
Dynamic Grows as tokens are generated; it is the default in Hugging Face Transformers. Does not require reserving the full maximum cache size at the outset. It supports sliding-window or chunked behavior when the model’s layers impose those limits. Changing cache shapes can make compilation less straightforward than with a fixed allocation. Actual compatibility and performance depend on the implementation.
Static Preallocates a cache up to a maximum size. A fixed shape can enable compilation and may be useful when predictable allocation is important. Can reserve memory for positions that are never used. Depending on implementation, attention may also do work on masked positions when a request is shorter than the allocation.
Quantized Stores cache values at reduced precision or in a compressed representation. Can reduce cache memory use, potentially allowing longer contexts or more concurrent sequences within a memory budget. Quality, latency, supported models, and compatibility depend on the quantization method and implementation; reduced precision is not a free or uniform improvement.
Offloaded Keeps most layer caches on CPU memory and transfers cache data as needed. Can relieve pressure on GPU memory. Transfers between CPU and GPU can lower throughput or increase latency. The effect depends on hardware, transfer costs, and workload.

When comparing implementations, measure the factors that matter for the workload: cache memory at the target context and batch size, decode latency or throughput, compilation support, sliding-window compatibility, precision effects, and complexity of integration. A configuration that suits one model or serving setup may be a poor fit for another.

Where cache optimization is still evolving

Standard cache strategies change how existing key and value states are stored or moved. Research can also change the architecture that produces those states. Cross-Layer Attention, presented at NeurIPS 2024, is an example: it shares key/value heads between adjacent layers to reduce KV-cache size. It is a research architecture, not a guaranteed drop-in cache setting for arbitrary existing models.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.