Temperature
Temperature is a number, usually between 0 and 2 depending on the provider. It controls how sharply a model favors its own most likely next token — a small chunk of text, often a word or part of a word, that the model generates one at a time. It doesn't change what the model knows or which tokens are possible. It changes how much the model's confidence gets exaggerated or flattened before a token is chosen.
Before picking the next token, a model scores every possible token, then converts those scores into a probability distribution — a list of every candidate token with a percentage chance attached to each one, all adding up to 100%. Temperature scales the scores before that conversion. Low temperature stretches the gap between the top candidate and everything else, so the most likely token wins almost every time. High temperature compresses that gap, so lower-probability tokens get a real chance.
What it looks like at different settings
Ask a model to complete "The capital of France is" at different temperatures:
- Low (near 0): "Paris." Same answer almost every time.
- Medium (around 1, most providers' default): "Paris." Still almost always correct, slightly more variation in phrasing.
- High (near the top of the range): occasionally still "Paris," but increasingly likely to wander into a plausible-sounding wrong city, or text that trails off-topic.
Factual-recall prompts don't show off temperature well because there's usually one dominant right answer. Open-ended prompts show it clearly — "write a tagline for a coffee shop" gives near-identical taglines at low temperature and a genuinely varied set at high temperature.
When low temperature is useful
Anything with a correct or preferred answer, where you want the model to commit to it: data extraction, classification, code generation, factual Q&A, calling tools with structured arguments. Consistency matters more than variety.
When higher temperature is useful
Anything where variety is the point: brainstorming, creative writing, generating multiple options to choose from, roleplay. Same output every time would defeat the purpose.
When to leave it alone
Don't use temperature to fix a different problem. Raising it to fix repetitive output in a data-extraction or tool-calling task just introduces wrong answers — those tasks want low temperature, period. Lowering it to fix a factually wrong answer won't help either: temperature changes phrasing and token variety, not whether the answer is correct. It's also not a substitute for reproducibility — see below.
Common misconception: temperature 0 does not guarantee identical output every time. Anthropic and OpenAI both document this: floating-point arithmetic on real hardware isn't perfectly reproducible, and some providers adjust behavior near 0 on their own. Temperature 0 is close to deterministic, not a guarantee of it. Byte-for-byte reproducibility needs more than a low temperature setting.
Practical recommendations
- Ranges and defaults differ by provider: Anthropic's Messages API runs 0–1, default 1. OpenAI's runs 0–2, also default 1. Check your provider's current docs rather than assuming a range.
- If output feels too erratic, lowering temperature is usually a more direct fix than rewriting the prompt.
Temperature vs. top-p
They're separate levers, not the same mechanism. Temperature scales the whole probability distribution — it changes how sharply the top candidate is favored. Top-p restricts the candidate pool itself. It keeps only the smallest set of tokens whose combined probability crosses a threshold, and it does that before temperature is even applied. Say the model's top candidates are "happy" (40%), "excited" (30%), "thrilled" (20%), and a long tail of unlikely words splitting the rest. A top-p of 0.9 keeps only "happy," "excited," and "thrilled," the smallest group that adds up past 90%. Everything else in that long tail gets dropped before temperature gets a say in anything. OpenAI's docs specifically recommend adjusting one or the other, not both at once — moving both makes output harder to reason about for little added benefit. Not every provider states this explicitly, but the reasoning holds generally.