Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

LLM serving is a coordination problem: the system must fit active requests’ growing key-value (KV) caches into accelerator memory while deciding which requests receive compute at each step. Cache management determines how much concurrent work can fit; scheduling determines how that work is processed without wasting capacity or compromising latency.

Why serving depends on both memory and scheduling

During autoregressive inference, a model generates output one token at a time. To avoid recomputing the entire context at every step, it retains key and value tensors from earlier tokens in a KV cache. Each active request needs its own cache, and that cache grows as the prompt and generated sequence grow.

Requests differ in prompt length and output length, so cache demand changes over time. If memory is fragmented or cache data is duplicated unnecessarily, less of the accelerator’s memory is available for useful requests. That limits how many sequences the server can keep active together. The PagedAttention paper identifies these problems as constraints on batching and serving capacity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Memory alone does not decide what happens next. At each model iteration, a serving system also chooses which active requests to process. It needs to admit work that fits available resources and form batches that use compute effectively while meeting latency goals. The cache policy and the scheduler therefore affect one another: fitting more requests can create more potential work, but the scheduler still has to decide how to serve it.

How a serving scheduler chooses work

Capacity: which requests can remain active?

A capacity decision asks whether new or continuing requests fit in the resources available, including KV-cache space. If a request cannot be supported, the server may need to defer it or manage active work differently. Capacity is dynamic because cached sequences grow during generation.

Microbatching: which requests run in the next step?

Once the system has selected work that can fit, it forms the batch for a forward pass. TensorRT-LLM’s PyTorch scheduler documentation describes distinct CapacityScheduler and MicroBatchScheduler roles: one stage considers capacity and another selects context and generation requests for microbatches. The guide is on the project’s main branch, so its behavior should be checked against the version actually deployed.

Why prompt prefill and token decode need different treatment

Serving has two different kinds of work. Prefill processes the input prompt, often handling many tokens together. Decode produces the response incrementally, typically one token per request per iteration. A long prefill can compete with ongoing decode work and make iteration times uneven, affecting latency for requests already generating.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sarathi-Serve addresses this scheduling tension with chunked prefill: it splits prompt processing into smaller pieces so new requests can join ongoing decode work without stalling those decodes, as described in the Sarathi-Serve paper. Chunking is a scheduling strategy, not a universal guarantee of better results; its value depends on the workload and latency objective.

How major designs manage cache and scheduling

Design Core idea What to examine
PagedAttention / vLLM Uses fixed-size KV blocks and block mapping, allowing dynamic allocation and cache sharing rather than requiring each sequence’s cache to occupy one contiguous physical region. The paper reports near-zero KV-cache waste as a system result. Cache capacity and sharing, kernel implementation, block-management overhead, throughput, and latency under matched workloads.
Sarathi-Serve Uses chunked prefills and stall-free schedules to balance prompt work with ongoing decode. Chunk size, prefill-to-decode mix, tail-latency target, hardware, parallelism, and serving capacity.
TensorRT-LLM scheduler Separates resource-capacity selection from microbatch selection at each step. Admission policy, KV-cache capacity, batch formation, paused requests, and workload behavior.
vAttention Reserves contiguous virtual address space while mapping physical memory on demand using CUDA virtual-memory mechanisms. Kernel compatibility, physical allocation granularity, runtime overhead, portability, and throughput under the paper’s test conditions.

These are different system design choices, not a product ranking. Compare them on the same model, accelerator, input and output lengths, concurrency, latency objective, and implementation version whenever possible.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What reported performance figures do—and do not—show

Published gains are conditional on the authors’ test setup. Sarathi-Serve’s authors reported 2.6× higher serving capacity for Mistral-7B on one A100 and up to 3.7× for Yi-34B on two A100 GPUs compared with vLLM. They also reported up to 5.6× end-to-end serving-capacity gain for Falcon-180B using pipeline parallelism. These are results from that paper’s evaluations, not general guarantees for other models, hardware, or traffic patterns.

The vAttention authors reported up to 1.23× serving throughput compared with PagedAttention-based FlashAttention and FlashInfer kernels in their evaluation. They also gave per-token KV-memory examples of 64 KB for Yi-6B, 128 KB for Llama-3-8B, and 240 KB for Yi-34B; those figures apply to the model configurations in the paper. The vAttention paper provides the relevant evaluation context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not combine these figures into a cross-paper leaderboard: the models, hardware, baselines, and methods differ. A useful comparison must state the workload and hardware, along with the concurrency and latency target; a headline multiplier without those conditions is not enough to predict performance in another deployment.

What to check when operating a serving system

  • Pin the implementation version. Defaults and feature availability change. The vLLM stable CLI reference documents KV-cache sizing and dtype controls, optional CPU KV-cache offloading, a scheduler admission watermark, asynchronous scheduling, and other serving options. It does not establish a best setting for every workload.
  • Track cache demand as sequences grow. Prompt length, generated length, and the number of active requests all influence how much cache capacity is needed.
  • Distinguish admission from batch formation. A request can fit the resource budget yet still not belong in a particular microbatch or iteration.
  • Evaluate the prefill/decode mix. A workload dominated by long prompts may stress the scheduler differently from one dominated by ongoing token generation.
  • Compare under matched conditions. Record model and configuration, accelerator count, parallelism, input and output lengths, concurrency, implementation version, and latency objective.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.