AI glossary · Models and context
What is top-p (nucleus sampling)?
Also called: top_p, nucleus sampling, topP
Definition
Top-p, or nucleus sampling, is a sampling setting that keeps only the smallest set of likeliest next tokens whose probabilities add up to at least p, then picks the next token from that set.
Explained
How it works
Sort the candidate tokens from most to least likely and add up their probabilities until the total reaches p. Those tokens are the nucleus. Everything after them is dropped, and the kept probabilities are rescaled to sum to 1 before one is sampled. When the model is confident the nucleus is one or two tokens; when it is unsure, it grows.
The method comes from the 2019 paper “The Curious Case of Neural Text Degeneration” (ICLR 2020). It found that always taking the likeliest words makes text bland and repetitive, while sampling from the full distribution lets in an unreliable long tail. Nucleus sampling cuts the tail and samples from what is left.
OpenAI’s reference calls top_p an alternative to sampling with temperature and recommends changing one or the other, not both. Newer models restrict it: OpenAI’s GPT-6 guide says to remove it when reasoning effort isn’t none, Anthropic accepts only 0.99 or above on models released after Claude Opus 4.6, and Google has deprecated top_p from Gemini 3.6 Flash onward, where the API ignores it.
Example
Top-p 0.9 and 0.5 on five candidates
Suppose the next-token probabilities are blue 0.50, green 0.25, red 0.12, purple 0.08, beige 0.05. With top_p 0.9, the running total passes 0.9 at the fourth token (0.95), so “beige” can never be chosen and the rest are rescaled by dividing by 0.95.
With top_p 0.5, “blue” reaches 0.50 on its own, so the nucleus is a single token and the model always picks it, just like greedy decoding.
| Token | Probability | Running total | Chance after top-p 0.9 |
|---|---|---|---|
| blue | 0.50 | 0.50 | 52.6% |
| green | 0.25 | 0.75 | 26.3% |
| red | 0.12 | 0.87 | 12.6% |
| purple | 0.08 | 0.95 | 8.4% |
| beige | 0.05 | 1.00 | Dropped |
Probabilities invented for the illustration; the cut and rescaling follow the paper’s definition.
Cost and quality
Why it matters
Cutting the tail removes rare, odd tokens that can send a sentence off course, without making every reply identical. Because the cut adapts to how confident the model is, it avoids the weakness of a fixed top-k, which keeps the same number of tokens whether the model is sure or unsure.
For most API work the provider’s default is the place to start, and on many current models you can’t change it at all. If you do tune, change one sampling setting at a time and measure on your own prompts.
Don’t mix up
Common confusions
- Top-p vs top-k
- Top-k keeps a fixed number of tokens, whatever their probabilities. Top-p keeps however many it takes to reach the probability share. Gemini’s API combines both on older models that support top-k, but deprecates both from Gemini 3.6 Flash onward. Anthropic has deprecated
top_kand rejects it on models released after Claude Opus 4.6. - Top-p vs temperature
- Temperature reshapes every probability and drops nothing. Top-p drops the tail and leaves the shape of what remains. Both change which token gets picked, so changing them together makes the effect hard to predict.
Go deeper
Try it and read more
Related
Related terms
- TemperatureTemperature is a sampling setting that divides a model’s next-token scores before they become probabilities, so low values make the likeliest token dominate and high values spread the choice across more tokens.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- HallucinationA hallucination is a confident, plausible-sounding statement from a language model that is false or unsupported by its input, such as an invented citation, date, function or quote.