What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Language model inference is the stage where a trained model is run on new input to produce output, such as the next words of a reply. Each time someone sends a prompt to a deployed model, inference is the computation that turns that prompt into a response. It is separate from training, the earlier process that adjusts the model’s parameters. In the common autoregressive text-generation path, the model first processes the whole prompt (prefill), then generates output one token at a time (decode), reusing cached attention state so it does not recompute earlier tokens at every step.
Inference, training and serving are different things
In machine learning, inference means executing a model after training is finished, using new inputs to compute predictions or generated outputs. Training is different: it changes the model’s parameters by learning from data. Inference leaves the parameters fixed.
Inference is often confused with serving. Inference is the model computation itself. Serving is the system around that computation: receiving requests, queueing them, batching them onto hardware, streaming partial output back to users, and recording metrics. A model can be inferred without any serving layer, for example in a single script on one machine. A production chatbot needs both.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →For a language model, input text is first converted into tokens, which are the units the model reads and writes. Tokens are often word fragments rather than whole words. A generative model then emits output tokens according to its generation method. The sequence described below is the widely used autoregressive, decoder-only path. Not every language model or generation architecture follows exactly these steps.
#1 Best Overall
What happens during a request
1. Tokenization and request setup
The prompt is converted into a sequence of tokens using the model’s tokenizer. This step matters for measurement as well as correctness. A token in one tokenizer may correspond to a different amount of text in another, so tokens per second from two models are not directly comparable unless the tokenizers are the same.
2. Prefill
During prefill, the model processes the entire input context in one pass and computes attention state for every prompt token. This is the step that reads the prompt. Its cost grows with prompt length, and it is the reason a long prompt can delay the first word of the answer.
3. Decode
Decode produces the output. An autoregressive model generates one token, appends it to the context, and then generates the next. Because each new token depends on the ones before it, the steps must run in sequence. Keys and values computed for earlier tokens are reused from the KV cache rather than recalculated at each step.
Rank #2
4. Stop and return
Generation continues until a stopping condition is met. Common conditions are a model-specific end-of-sequence token or a configured maximum output length. These rules are deployment and application settings, not universal constants. A serving system may stream tokens as they are produced, or it may wait and return the complete response.
Why the KV cache matters
The KV cache stores the attention keys and values for tokens the model has already processed. Without it, every decode step would have to recompute attention information for the whole preceding sequence. With it, each step adds only the new token’s entries. This is what makes sequential generation practical.
The cache is not free. Its memory footprint depends on the model architecture, numerical precision, sequence length, and the number of requests active at once. Long contexts and high concurrency can exhaust accelerator memory even when the model weights fit comfortably. A reasonable one-line description is that the cache saves repeated work by keeping the model’s attention state for earlier tokens, and that state takes memory. It does not remove all repeated computation, and it does not improve performance in every setting regardless of memory limits.
Batching, colocation and other trade-offs
Batching
Batching lets hardware process several requests together, which can raise utilization and total output per unit of time. Static batching groups requests and runs them as a unit, so a short request can wait for a longer one in the same batch. Continuous, or in-flight, batching lets the serving engine add and remove requests as work progresses. Whether batching helps depends on arrival patterns, prompt and output lengths, model size, hardware, and latency targets. Higher aggregate throughput can come at the cost of slower responses for individual users.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Colocated and disaggregated serving
In colocated serving, prefill and decode run on the same GPU resources. NVIDIA’s TensorRT-LLM documentation notes that prefill work can interfere with token generation, which affects token-to-token latency. Disaggregated serving assigns the two phases to separate GPU pools so each can be tuned independently. The cost is that the KV-cache blocks produced during prefill must be transferred to the decode pool. NVIDIA’s documentation describes workloads with long inputs and moderate outputs as a case where separation can help. That is a workload-specific observation, not a general recommendation.
Quantization and parallelism
Quantization stores weights or performs computation at lower numerical precision. It can reduce memory use and serving cost, but it can also change output quality, and the effect varies by model and hardware. It should be tested on the actual model and task. Model parallelism splits a model across several accelerators when it does not fit on one. It adds communication overhead and operational complexity.
Rank #4
The four measures used to describe inference speed
Speed claims for inference use several different measurements. They describe different parts of the experience, so a headline number is meaningful only when its definition is known. The table below follows the definitions in NVIDIA’s benchmarking documentation.
| Measure | What it captures | What to check |
|---|---|---|
| Time to first token (TTFT) | Time from query submission to the first output token received | Usually includes queueing, prefill and network latency, so longer prompts raise it |
| End-to-end request latency | Time from submission until the full response arrives | Includes queueing, batching and network latency along with generation |
| Inter-token latency (ITL), also called time per output token (TPOT) | Average time between successive output tokens | Tools differ on whether TTFT is included in the average; NVIDIA’s AIPerf definition excludes it |
| Tokens per second (TPS) | Output token rate, either aggregate across the system or per request | Confirm which definition is reported; aggregate figures can rise with concurrency while per-user speed falls |
How to read a claimed inference speed
A single figure such as “fast inference” or “X tokens per second” is incomplete on its own. Before comparing two results, check that the following are stated:
- The model name and version, and the tokenizer used to count tokens.
- The prompt and output length distributions, not only an average.
- The request arrival rate and concurrency level.
- The decoding settings that affect output length and computation.
- The hardware, including accelerator type and count.
- The serving software and version, since optimizations change between releases.
- The exact formula for each metric, including whether TTFT is included in ITL.
No single inference speed applies across models, hardware and workloads. Published benchmark tables are useful only for the configurations they describe. The numerical examples in NVIDIA’s technical writing on inference optimization illustrate calculation methods, such as memory estimates for assumed model settings, and should not be read as typical measured results.
Best Value
Documentation for inference tools changes with each software release. Confirm the current behaviour of a specific engine before relying on a particular default or metric definition.
When you need a reference point, use the vendor’s documentation for the software you are running, and test with a workload that resembles your own.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

