Temperature and Top-p in LLM
Temperature and top-p are LLM sampling parameters that control how a large language model picks each next token from its probability scores. Temperature makes the scores sharper or flatter, and top-p removes unlikely tokens, so together they decide whether output is predictable or varied.
- Sampling: It is the decoding step in which a large language model selects one token from a list of scored candidates.
- Temperature: It divides the raw scores before normalisation, so low values concentrate probability on the most likely token.
- Top-p: It restricts sampling to the most probable tokens whose cumulative probability reaches a threshold, such as 0.9.
- Logits: They are the unnormalised scores the model calculates for every token in its vocabulary.
- Request parameters: Most commercial LLM APIs accept
temperatureandtop_pas fields in each request.
For example, after for i in , a code assistant at low temperature almost always suggests range(, while a higher temperature sometimes suggests enumerate( or zip(.
Key Characteristics of Temperature and Top-p
- Temperature scale: A value of 1.0 leaves the distribution unchanged, lower values sharpen it, and higher values flatten it.
- Greedy decoding: A temperature near 0 approximates greedy decoding, which always selects the single most probable token.
- Adaptive cut-off: Top-p keeps few candidates when the model is confident and more candidates when the probabilities are similar.
- Renormalisation: After top-p removes candidates, the remaining probabilities are rescaled so that their total is exactly 1 again.
- Not a quality control: Neither setting adds knowledge or reasoning, so a higher temperature produces variety rather than accuracy.
- Provider defaults: Default values and permitted ranges differ between providers, models and API versions.
How Temperature and Top-p Work
- Scoring: The model produces a logit for every token in its vocabulary, based on the text so far. The full walkthrough is in how LLMs work.
- Temperature scaling: Each logit is divided by the temperature value, which widens or compresses the differences between competing scores.
- Softmax: The scaled logits are converted into probabilities by the softmax function, which makes every value positive and the total equal to 1.
- Top-p filtering: Candidates are sorted by probability, and the smallest leading group whose cumulative total reaches the top-p value is retained.
- Sampling: One token is drawn randomly from the retained candidates, weighted by their renormalised probabilities.
- Repetition: The selected token is appended to the input sequence, and the process repeats for the following position until the response is complete.
Temperature Scaling and Nucleus Sampling
Temperature is a single division inside the softmax function, and top-p is a selection rule that keeps the smallest group of probable tokens. Both operations act on the same list of logits produced by the model.
Softmax Temperature Formula
- : the logit, meaning the unnormalised score of token .
- : the temperature, a positive number such as 0.5, 1 or 2.
- : the probability of token after temperature scaling.
The effect is easiest to understand through the ratio between two tokens, which depends only on the difference between their logits divided by the temperature.
A low temperature enlarges every difference, so the leading token dominates the distribution. A high temperature compresses every difference, so the probabilities move towards equal values. As the temperature approaches zero, sampling becomes equivalent to greedy decoding.
The Top-p (Nucleus) Set
- : the top-p threshold, commonly a value such as 0.9.
- : the probability of a token after temperature scaling.
- : the nucleus, constructed by adding tokens in descending order of probability until their cumulative total reaches the threshold.
- : the renormalised probability that is used for the final random selection.
Worked Example at T = 0.5, 1 and 2
The logits of range( and enumerate( are 4.0 and 3.2, so their difference is 0.8 before temperature scaling.
- Low temperature: At 0.5,
range(becomes almost five times as probable asenumerate(, so completions become highly consistent. - Neutral temperature: At 1, the original distribution is preserved, and the leading token is a little more than twice as probable.
- High temperature: At 2, the ratio falls below one and a half, so alternative completions appear much more frequently.
- Nucleus construction: At temperature 1, the running total first reaches the threshold after three tokens, so the nucleus contains three candidates.
The program below calculates all five probabilities at each temperature and the nucleus size for a threshold of 0.9.
import math
# Logits for the token after "for i in " (illustrative, as above)
logits = {"range(": 4.0, "enumerate(": 3.2, "items": 2.5, "zip(": 1.4, "reversed(": 0.6}
def softmax_t(scores, t):
exps = {tok: math.exp(z / t) for tok, z in scores.items()}
total = sum(exps.values())
return {tok: e / total for tok, e in exps.items()}
def nucleus(probs, p):
# Smallest set of top tokens whose probabilities add up to at least p
kept, running = [], 0.0
for tok, pr in sorted(probs.items(), key=lambda x: -x[1]):
kept.append(tok)
running += pr
if running >= p:
return kept, running
print("token T=0.5 T=1.0 T=2.0")
table = {t: softmax_t(logits, t) for t in (0.5, 1.0, 2.0)}
for tok in logits:
print(f"{tok:11}", " ".join(f"{table[t][tok]:.3f}" for t in table))
for t, probs in table.items():
kept, mass = nucleus(probs, 0.9)
print(f"T={t}: top-p 0.9 keeps {len(kept)} tokens (mass {mass:.3f})")token T=0.5 T=1.0 T=2.0
range( 0.795 0.562 0.385
enumerate( 0.160 0.252 0.258
items 0.040 0.125 0.182
zip( 0.004 0.042 0.105
reversed( 0.001 0.019 0.070
T=0.5: top-p 0.9 keeps 2 tokens (mass 0.955)
T=1.0: top-p 0.9 keeps 3 tokens (mass 0.940)
T=2.0: top-p 0.9 keeps 4 tokens (mass 0.930)- Ratios confirm the formula: At temperature 0.5, the printed probabilities give a ratio of approximately 4.97, which matches the exponential calculation within rounding.
- Nucleus size follows temperature: The same threshold keeps two tokens at temperature 0.5 but four tokens at temperature 2, because a flatter distribution requires more candidates to reach the same cumulative probability.
Because temperature also changes how many tokens top-p retains, adjusting one parameter at a time makes its effect on code completions considerably easier to measure.
Example: Temperature Scaling and Top-p Filtering in Python
The program below applies both settings to five illustrative scores for the token after for i in .
import math
# Raw scores (logits) for the token after "for i in " (illustrative)
logits = {"range(": 4.0, "enumerate(": 3.2, "items": 2.5, "zip(": 1.4, "reversed(": 0.6}
def softmax(scores, temperature):
scaled = {t: s / temperature for t, s in scores.items()}
top = max(scaled.values())
exps = {t: math.exp(s - top) for t, s in scaled.items()}
total = sum(exps.values())
return {t: e / total for t, e in exps.items()}
def top_p(probs, p):
# Keep the smallest set of top tokens whose probabilities add up to p
kept, running = {}, 0.0
for t, pr in sorted(probs.items(), key=lambda x: -x[1]):
kept[t] = pr
running += pr
if running >= p:
break
total = sum(kept.values())
return {t: pr / total for t, pr in kept.items()}
def show(label, probs):
print(label, ", ".join(f"{t} {pr:.2f}" for t, pr in probs.items()))
for temp in (0.5, 1.0, 1.5):
show(f"T={temp}:", softmax(logits, temp))
show("T=1.0, top_p=0.9:", top_p(softmax(logits, 1.0), 0.9))T=0.5: range( 0.79, enumerate( 0.16, items 0.04, zip( 0.00, reversed( 0.00
T=1.0: range( 0.56, enumerate( 0.25, items 0.13, zip( 0.04, reversed( 0.02
T=1.5: range( 0.45, enumerate( 0.26, items 0.16, zip( 0.08, reversed( 0.05
T=1.0, top_p=0.9: range( 0.60, enumerate( 0.27, items 0.13- Low temperature: At 0.5,
range(rises to 0.79 and the two weakest tokens fall to almost zero, so completions become consistent. - High temperature: At 1.5, the weakest token
reversed(rises from 0.02 to 0.05, so unusual completions appear more often. - Top-p cut: With top-p 0.9, the first two tokens reach only 0.81, so
itemsis also kept, whilezip(andreversed(are removed.
Applications of Temperature and Top-p
- Code completion: Low temperature keeps inline suggestions stable and consistent with conventional coding patterns.
- Structured output: Low values reduce formatting errors when a model must return structured JSON output.
- Test data generation: Higher values produce more diverse sample inputs, identifiers and edge cases.
- Brainstorming: Higher values generate a wider variety of variable names, commit messages or architectural alternatives.
- Multiple candidates: Systems that generate several solutions and select one require some randomness to produce distinct candidates.
- Evaluation runs: Fixed low values make repeated evaluation runs easier to compare, as explained in AI agent evaluation.
Advantages
- Simple configuration: Two numeric parameters change output behaviour without fine-tuning or retraining the model.
- Per-request tuning: Each API call can use different values for different tasks within the same application.
- Fewer anomalous tokens: Top-p removes the long tail of improbable tokens that produce irrelevant or incoherent output.
- Predictability on demand: Low temperature supports repeatable responses in automated pipelines and continuous integration checks.
Limitations
- No factual guarantee: A temperature of 0 still produces confident errors, a problem covered in LLM hallucination.
- Not fully deterministic: Identical requests can still differ slightly because of hardware and batching effects.
- Interacting settings: Changing both settings at once makes results hard to reason about, so providers often advise tuning one.
- Model-specific behaviour: The same value can behave differently across models, so tuned settings rarely transfer between them.
- Restricted on some models: Some reasoning models ignore these parameters or fix them at predetermined values.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does a temperature below 1.0 do to next-token probabilities?
Frequently Asked Questions
What is a good temperature for code generation?
Code generation usually works well with a low temperature, because correct code follows common patterns. Higher values are useful only when several different candidate solutions are needed.
Temperature vs top-p: what is the difference?
Temperature changes the shape of the whole probability distribution before sampling. Top-p removes the unlikely tokens and samples only from the most likely group that covers a set share of probability.
Does temperature 0 make an LLM deterministic?
Temperature 0 makes the model pick the most likely token at each step, so outputs become very consistent. Small differences can still appear because of how requests are batched and computed on hardware.
Is top-p the same as nucleus sampling?
Yes, top-p sampling is also called nucleus sampling. The nucleus is the smallest group of top tokens whose probabilities add up to the chosen value.
Does a higher temperature make an LLM more creative?
A higher temperature makes less likely tokens more common, so output becomes more varied. It does not add knowledge or reasoning, and very high values often produce incoherent text.
Related Articles
- How LLMs WorkLearn how LLMs work from training to inference: tokenization, embeddings, attention layers and decoding, with a Python trace of one forward pass.
- Tokens and Tokenization in LLMLearn how tokenization in LLMs splits text and code into subword tokens and why token counts drive cost and context, with a Python tokenizer example.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- Structured Output (JSON) from LLMsLearn how structured output gets JSON from an LLM: schemas, extraction prompts, constrained decoding and validation, with a Python bug report JSON check.
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.
- Transformer Architecture ExplainedTransformer architecture explained: self-attention, query, key and value vectors, multi-head attention and stacked layers, with Python attention code.