What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Language model inference is the stage where a trained model is run on new input to produce output, such as the next words of a reply. Each time someone sends a prompt to a deployed model, inference is the computation that turns that prompt into a response. It is separate from training, the earlier process that adjusts the model’s parameters. In the common autoregressive text-generation path, the model first processes the whole prompt (prefill), then generates output one token at a time (decode), reusing cached attention state so it does not recompute earlier tokens at every step.

Inference, training and serving are different things

In machine learning, inference means executing a model after training is finished, using new inputs to compute predictions or generated outputs. Training is different: it changes the model’s parameters by learning from data. Inference leaves the parameters fixed.

Inference is often confused with serving. Inference is the model computation itself. Serving is the system around that computation: receiving requests, queueing them, batching them onto hardware, streaming partial output back to users, and recording metrics. A model can be inferred without any serving layer, for example in a single script on one machine. A production chatbot needs both.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a language model, input text is first converted into tokens, which are the units the model reads and writes. Tokens are often word fragments rather than whole words. A generative model then emits output tokens according to its generation method. The sequence described below is the widely used autoregressive, decoder-only path. Not every language model or generation architecture follows exactly these steps.

What happens during a request

1. Tokenization and request setup

The prompt is converted into a sequence of tokens using the model’s tokenizer. This step matters for measurement as well as correctness. A token in one tokenizer may correspond to a different amount of text in another, so tokens per second from two models are not directly comparable unless the tokenizers are the same.

2. Prefill

During prefill, the model processes the entire input context in one pass and computes attention state for every prompt token. This is the step that reads the prompt. Its cost grows with prompt length, and it is the reason a long prompt can delay the first word of the answer.

3. Decode

Decode produces the output. An autoregressive model generates one token, appends it to the context, and then generates the next. Because each new token depends on the ones before it, the steps must run in sequence. Keys and values computed for earlier tokens are reused from the KV cache rather than recalculated at each step.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Stop and return

Generation continues until a stopping condition is met. Common conditions are a model-specific end-of-sequence token or a configured maximum output length. These rules are deployment and application settings, not universal constants. A serving system may stream tokens as they are produced, or it may wait and return the complete response.

Why the KV cache matters

The KV cache stores the attention keys and values for tokens the model has already processed. Without it, every decode step would have to recompute attention information for the whole preceding sequence. With it, each step adds only the new token’s entries. This is what makes sequential generation practical.

The cache is not free. Its memory footprint depends on the model architecture, numerical precision, sequence length, and the number of requests active at once. Long contexts and high concurrency can exhaust accelerator memory even when the model weights fit comfortably. A reasonable one-line description is that the cache saves repeated work by keeping the model’s attention state for earlier tokens, and that state takes memory. It does not remove all repeated computation, and it does not improve performance in every setting regardless of memory limits.

Batching, colocation and other trade-offs

Batching

Batching lets hardware process several requests together, which can raise utilization and total output per unit of time. Static batching groups requests and runs them as a unit, so a short request can wait for a longer one in the same batch. Continuous, or in-flight, batching lets the serving engine add and remove requests as work progresses. Whether batching helps depends on arrival patterns, prompt and output lengths, model size, hardware, and latency targets. Higher aggregate throughput can come at the cost of slower responses for individual users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Colocated and disaggregated serving

In colocated serving, prefill and decode run on the same GPU resources. NVIDIA’s TensorRT-LLM documentation notes that prefill work can interfere with token generation, which affects token-to-token latency. Disaggregated serving assigns the two phases to separate GPU pools so each can be tuned independently. The cost is that the KV-cache blocks produced during prefill must be transferred to the decode pool. NVIDIA’s documentation describes workloads with long inputs and moderate outputs as a case where separation can help. That is a workload-specific observation, not a general recommendation.

Quantization and parallelism

Quantization stores weights or performs computation at lower numerical precision. It can reduce memory use and serving cost, but it can also change output quality, and the effect varies by model and hardware. It should be tested on the actual model and task. Model parallelism splits a model across several accelerators when it does not fit on one. It adds communication overhead and operational complexity.

The four measures used to describe inference speed

Speed claims for inference use several different measurements. They describe different parts of the experience, so a headline number is meaningful only when its definition is known. The table below follows the definitions in NVIDIA’s benchmarking documentation.

Measure What it captures What to check
Time to first token (TTFT) Time from query submission to the first output token received Usually includes queueing, prefill and network latency, so longer prompts raise it
End-to-end request latency Time from submission until the full response arrives Includes queueing, batching and network latency along with generation
Inter-token latency (ITL), also called time per output token (TPOT) Average time between successive output tokens Tools differ on whether TTFT is included in the average; NVIDIA’s AIPerf definition excludes it
Tokens per second (TPS) Output token rate, either aggregate across the system or per request Confirm which definition is reported; aggregate figures can rise with concurrency while per-user speed falls
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to read a claimed inference speed

A single figure such as “fast inference” or “X tokens per second” is incomplete on its own. Before comparing two results, check that the following are stated:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The model name and version, and the tokenizer used to count tokens.
  • The prompt and output length distributions, not only an average.
  • The request arrival rate and concurrency level.
  • The decoding settings that affect output length and computation.
  • The hardware, including accelerator type and count.
  • The serving software and version, since optimizations change between releases.
  • The exact formula for each metric, including whether TTFT is included in ITL.

No single inference speed applies across models, hardware and workloads. Published benchmark tables are useful only for the configurations they describe. The numerical examples in NVIDIA’s technical writing on inference optimization illustrate calculation methods, such as memory estimates for assumed model settings, and should not be read as typical measured results.

Documentation for inference tools changes with each software release. Confirm the current behaviour of a specific engine before relying on a particular default or metric definition.

When you need a reference point, use the vendor’s documentation for the software you are running, and test with a workload that resembles your own.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.