Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Top-p, also called nucleus sampling, limits the next-token choices a language model can sample from by keeping the smallest group of most-probable tokens whose probabilities add up to a chosen threshold. The number of eligible tokens can change at every generation step because the model’s probability distribution changes. Unlike top-k, which keeps a fixed number of candidates, top-p keeps a variable-size pool.

How top-p sampling works

At each step of text generation, a language model assigns a probability to each possible next token. Top-p sorts those tokens from most to least probable, then keeps the smallest prefix whose cumulative probability reaches the selected threshold, p. The retained probabilities are renormalized, and the model samples the next token from that pool. The process repeats for each new token.

For example, suppose the leading token probabilities are 0.30, 0.20, and 0.10, and the threshold is 0.50. The first two tokens together reach 0.50, so they are retained and the third is excluded. This is an instructional example from Google Cloud’s documentation, not a generally recommended setting.

Top-p does not mean “the top p percent of tokens.” It is a cutoff on cumulative probability mass. If likely tokens dominate the distribution, only a few may be needed to reach the threshold; if probability is spread across more candidates, the pool can be larger. As the distribution changes during generation, so can the pool size.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Top-p vs. top-k and temperature

Control What it changes Candidate pool
Top-p The cumulative probability mass eligible for sampling Variable in size; responds to the distribution at that generation step
Top-k The number of eligible tokens Fixed at k, even when the distribution is more or less concentrated
Temperature The probability distribution used for sampling, affecting randomness Does not itself specify a cumulative cutoff or fixed candidate count

A fixed top-k can cover a small or large share of the probability mass depending on how concentrated the next-token distribution is. Top-p instead targets a chosen share of that mass, allowing the candidate count to expand or contract.

Temperature is a separate control: it changes the distribution from which sampling occurs, while top-p filters candidates using cumulative probability. Their interaction and processing order depend on the model or runtime. For instance, NVIDIA’s TensorRT-Model-Connect documentation describes an implementation that applies temperature before softmax and top-p filtering. Some systems let you use top-k and top-p together, so check the relevant documentation for whether both are supported and which is applied first.

Why nucleus sampling was proposed

In “The Curious Case of Neural Text Degeneration,” published in 2019, Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi describe a tension in text generation: likelihood-oriented decoding can produce bland or repetitive text, while unrestricted sampling can draw from a long tail of low-probability tokens. Their proposed dynamic nucleus aims to truncate that less reliable tail while retaining diversity. They write that sampling from the dynamic nucleus “allows for diversity while effectively truncating the less reliable tail of the distribution.” Read the paper.

That rationale is not a guarantee that top-p will improve every prompt or task. Hugging Face’s maintained guide notes that there is no one-size-fits-all decoding method and that top-p and top-k can still produce repetition: How to generate text with Transformers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose and test a top-p value

There is no universal optimal top-p value established across models. As one illustration of how a threshold behaves, Hugging Face’s guide uses 0.92 and shows it retaining nine tokens for one distribution and three for another. That example demonstrates a changing pool size; it is not a recommendation for every model.

Google Cloud’s documentation advises that, within its documented setting, a lower top-P value produces less-random responses and a higher value produces more-random responses. Available parameters and their effects can differ by model, so treat that direction as platform-specific guidance rather than a universal prescription. Check the documentation for the exact model or API you use.

  1. Keep the model, prompt, and other generation settings fixed so the comparison is meaningful.
  2. Change top-p alone, using values supported by that model or API.
  3. Generate multiple samples at each setting; a single response may not show the range of possible outputs.
  4. Compare samples against the task’s actual criteria, such as accuracy, consistency, diversity, or suitability of tone.
  5. If you also adjust temperature or top-k, change one control at a time and check the runtime’s documented parameter order.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.