Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is temperature in an LLM?

Also called: sampling temperature, LLM temperature

Definition

Temperature is a sampling setting that divides a model’s next-token scores before they become probabilities, so low values make the likeliest token dominate and high values spread the choice across more tokens.

Explained

How it works

At each step a model gives every possible next token a score (a logit). Softmax turns the scores into probabilities, and one token is sampled from them. Temperature T divides every score first: each probability is exp(score ÷ T) divided by the sum of exp(score ÷ T) over all tokens. Below 1, the gaps widen and the favourite wins more often; above 1, they shrink and unlikely tokens get picked more.

Ranges differ: OpenAI and Gemini accept 0 to 2, Anthropic 0 to 1 with a default of 1. Newer models restrict it. OpenAI’s GPT-6 guide says to remove temperature when reasoning effort isn’t none. Anthropic marks it deprecated for models released after Claude Opus 4.6, which reject any value except 1.0. Google has deprecated temperature, top_p and top_k from Gemini 3.6 Flash and Gemini 3.5 Flash-Lite onward: the API ignores them now and will reject them in future model generations.

Example

Four candidate tokens at three temperatures

Say a model is completing “My favourite colour is” and gives four candidates these made-up scores: blue 2.0, green 1.0, red 0.5, purple −0.5. Real models score their whole vocabulary, but the arithmetic is the same.

At 1.0 the scores are used as they are and “blue” gets 59.8%. At 0.5 it rises to 83.9% and “purple” almost disappears (0.6%). At 1.5 “blue” falls to 48.3% and “purple” climbs to 9.1%. No token is ever removed; temperature only reweights. Cutting off the unlikely tail is what top-p does.

Probability of each token, by temperature
Token (score)T = 0.5T = 1.0T = 1.5
blue (2.0)83.9%59.8%48.3%
green (1.0)11.4%22.0%24.8%
red (0.5)4.2%13.3%17.8%
purple (−0.5)0.6%4.9%9.1%

Computed with softmax(score ÷ T). The scores are invented for the illustration.

Cost and quality

Why it matters

Low values suit extraction, classification and code, where you want the likeliest answer; higher values suit brainstorming. But on current reasoning models the provider has usually decided for you, so check the model’s docs before tuning: a value the API rejects returns an error, and newer Gemini models simply ignore it.

Temperature is not a reproducibility switch. Anthropic’s reference says that even at 0.0 results are not fully deterministic. For dependable output, use structured outputs, a pinned model snapshot and tests.

Don’t mix up

Common confusions

Temperature vs top-p
Temperature reshapes the whole distribution. Top-p cuts it down to the likeliest tokens that cover a share of the probability. OpenAI recommends changing one or the other, not both.
Low temperature vs accurate answers
A lower temperature makes the model more consistent, not more correct. It can give the same wrong answer every time; see hallucination.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary