AI glossary · Models and context
What is temperature in an LLM?
Also called: sampling temperature, LLM temperature
Definition
Temperature is a sampling setting that divides a model’s next-token scores before they become probabilities, so low values make the likeliest token dominate and high values spread the choice across more tokens.
Explained
How it works
At each step a model gives every possible next token a score (a logit). Softmax turns the scores into probabilities, and one token is sampled from them. Temperature T divides every score first: each probability is exp(score ÷ T) divided by the sum of exp(score ÷ T) over all tokens. Below 1, the gaps widen and the favourite wins more often; above 1, they shrink and unlikely tokens get picked more.
Ranges differ: OpenAI and Gemini accept 0 to 2, Anthropic 0 to 1 with a default of 1. Newer models restrict it. OpenAI’s GPT-6 guide says to remove temperature when reasoning effort isn’t none. Anthropic marks it deprecated for models released after Claude Opus 4.6, which reject any value except 1.0. Google has deprecated temperature, top_p and top_k from Gemini 3.6 Flash and Gemini 3.5 Flash-Lite onward: the API ignores them now and will reject them in future model generations.
Example
Four candidate tokens at three temperatures
Say a model is completing “My favourite colour is” and gives four candidates these made-up scores: blue 2.0, green 1.0, red 0.5, purple −0.5. Real models score their whole vocabulary, but the arithmetic is the same.
At 1.0 the scores are used as they are and “blue” gets 59.8%. At 0.5 it rises to 83.9% and “purple” almost disappears (0.6%). At 1.5 “blue” falls to 48.3% and “purple” climbs to 9.1%. No token is ever removed; temperature only reweights. Cutting off the unlikely tail is what top-p does.
| Token (score) | T = 0.5 | T = 1.0 | T = 1.5 |
|---|---|---|---|
| blue (2.0) | 83.9% | 59.8% | 48.3% |
| green (1.0) | 11.4% | 22.0% | 24.8% |
| red (0.5) | 4.2% | 13.3% | 17.8% |
| purple (−0.5) | 0.6% | 4.9% | 9.1% |
Computed with softmax(score ÷ T). The scores are invented for the illustration.
Cost and quality
Why it matters
Low values suit extraction, classification and code, where you want the likeliest answer; higher values suit brainstorming. But on current reasoning models the provider has usually decided for you, so check the model’s docs before tuning: a value the API rejects returns an error, and newer Gemini models simply ignore it.
Temperature is not a reproducibility switch. Anthropic’s reference says that even at 0.0 results are not fully deterministic. For dependable output, use structured outputs, a pinned model snapshot and tests.
Don’t mix up
Common confusions
- Temperature vs top-p
- Temperature reshapes the whole distribution. Top-p cuts it down to the likeliest tokens that cover a share of the probability. OpenAI recommends changing one or the other, not both.
- Low temperature vs accurate answers
- A lower temperature makes the model more consistent, not more correct. It can give the same wrong answer every time; see hallucination.
Go deeper
Try it and read more
Related
Related terms
- Top-pTop-p, or nucleus sampling, is a sampling setting that keeps only the smallest set of likeliest next tokens whose probabilities add up to at least p, then picks the next token from that set.
- HallucinationA hallucination is a confident, plausible-sounding statement from a language model that is false or unsupported by its input, such as an invented citation, date, function or quote.
- Structured outputsStructured outputs is an LLM API feature that constrains the model’s reply to a JSON Schema you supply, so the response parses and has the fields and types you asked for, unlike JSON mode, which only promises valid JSON.
- Model snapshotA model snapshot is a fixed version of a model behind a specific API model ID, so requests to that ID keep getting the same model until it is retired, unlike an alias that the provider can move.