Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
When an LLM generates a token, it uses the prompt and any other context supplied to it to score possible next tokens, selects one according to a decoding method, and adds that token to the sequence. It then repeats the process until a stopping condition is met. So, what happens when an LLM generates a token? And how does an LLM predict the next word? The key distinction is that the model predicts a token—not necessarily a whole word.
What counts as a token?
A token is a unit of text defined by a model’s tokenizer. Depending on that tokenizer, a token can represent a complete word, part of a word, punctuation, text associated with whitespace, or a special token. There is no universal rule that one token equals one word; token boundaries differ among models.
The model processes token IDs rather than directly choosing a finished word from a sentence. An application may also format the prompt with a chat template or include other context, so the visible user message is not necessarily the entire input the model receives. Hugging Face’s Transformers generation tutorial shows tokenizer-produced input IDs being passed to a model.
What happens during one token-generation step?
- Build the context. The prompt is represented in the format and token IDs expected by the model. This context provides the information the model uses to predict what comes next.
- Compute next-token scores. The model processes the context and produces logits—scores for possible tokens at the next position. A logit is not a word or a completed answer; it is an input to the selection process.
- Select a token. A decoding method uses those scores to choose a token. For example, greedy decoding picks the highest-scoring candidate, while sampling selects according to a probability distribution. Settings such as temperature can affect sampling when it is enabled.
- Append the token. The selected token ID is added to the sequence. The new sequence now provides the context for the next step.
- Repeat or stop. The model continues this loop until it produces an end-of-sequence token, reaches a maximum-new-token limit, or meets another configured stopping condition.
In the common autoregressive workflow described in the Hugging Face Transformers tutorial, one token is one iteration of the decoding loop—not the entire response. The model does not usually decide the whole final sentence in a single step.
#1 Best Overall
How do decoding methods choose the next token?
Decoding methods differ in how they turn candidate scores into a continuation. The choice can affect variation between runs and whether the process favors a locally likely next token or compares longer candidate sequences.
| Method | Selection rule | Variation and typical use |
|---|---|---|
| Greedy decoding | Selects the highest-scoring next token at each step. | Does not sample among alternatives at a given step; it favors the locally most likely choice. |
| Sampling | Selects from a probability distribution over possible tokens. | Can produce more varied continuations. Settings such as temperature affect selection when sampling is enabled. |
| Beam search | Tracks multiple candidate sequences and compares them by their overall probability. | Can be useful for input-grounded tasks; it is not simply another name for sampling or greedy decoding. |
These approaches are different selection strategies, not different definitions of a token. None is universally best: the appropriate method depends on the task and the desired behavior. Hugging Face outlines these decoding approaches in its generation strategies documentation.
Why does an LLM use a KV cache?
During generation, attention layers compute key and value representations for tokens in the context. A KV cache retains previously computed key and value states so later steps can reuse them instead of recalculating those states for the entire existing sequence. Once the prompt has populated the cache, a cached decoding step can process the newly added token while keeping prior states available.
This reuse can reduce repeated computation and speed up inference, but the retained states use memory, which grows with the context. The actual speed and memory trade-off depends on the model and runtime. Hugging Face’s KV cache guide describes its role in autoregressive generation. Cached and uncached execution should not be assumed to produce bit-for-bit identical output: differences in matrix-multiplication kernels can lead to slight numerical variation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What this explanation does—and does not—cover
This is the common autoregressive transformer decoding pattern documented for Hugging Face Transformers, not a guarantee that every LLM service uses the same software or performs precisely one model operation for each visible text token. Architectures, tokenizers, serving systems, stopping rules, and numerical behavior can vary; some systems also use speculative or multi-token decoding methods. The cited documentation explains software behavior but does not establish a universal measured time, compute requirement, or energy cost for generating one token.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

