…
Skip to content
Topics
On this page

Temperature and Top-p in LLM

Temperature and top-p are LLM sampling parameters that control how a large language model picks each next token from its probability scores. Temperature makes the scores sharper or flatter, and top-p removes unlikely tokens, so together they decide whether output is predictable or varied.

  • Sampling: It is the decoding step in which a large language model selects one token from a list of scored candidates.
  • Temperature: It divides the raw scores before normalisation, so low values concentrate probability on the most likely token.
  • Top-p: It restricts sampling to the most probable tokens whose cumulative probability reaches a threshold, such as 0.9.
  • Logits: They are the unnormalised scores the model calculates for every token in its vocabulary.
  • Request parameters: Most commercial LLM APIs accept temperature and top_p as fields in each request.
How temperature and top-p choose the next token in a code completionAfter the code "for i in", the model produces raw scores for five candidate tokens. The scores are divided by the temperature and turned into probabilities: range( 56 percent, enumerate( 25 percent, items 13 percent, zip( 4 percent and reversed( 2 percent at temperature 1.0. Top-p 0.9 keeps the first three tokens, which add up to 94 percent, and removes the last two. One token is then sampled from the kept group. The numbers are illustrative, not from a real model.for i in ...Scores ÷ temperatureProbabilities at T = 1.0 (illustrative)range( 56%enumerate( 25%items 13%zip( 4%reversed( 2%top-p 0.9 cutremoved belowSample one
How temperature and top-p choose the next token in a code completion

For example, after for i in , a code assistant at low temperature almost always suggests range(, while a higher temperature sometimes suggests enumerate( or zip(.

Key Characteristics of Temperature and Top-p

  • Temperature scale: A value of 1.0 leaves the distribution unchanged, lower values sharpen it, and higher values flatten it.
  • Greedy decoding: A temperature near 0 approximates greedy decoding, which always selects the single most probable token.
  • Adaptive cut-off: Top-p keeps few candidates when the model is confident and more candidates when the probabilities are similar.
  • Renormalisation: After top-p removes candidates, the remaining probabilities are rescaled so that their total is exactly 1 again.
  • Not a quality control: Neither setting adds knowledge or reasoning, so a higher temperature produces variety rather than accuracy.
  • Provider defaults: Default values and permitted ranges differ between providers, models and API versions.

How Temperature and Top-p Work

  1. Scoring: The model produces a logit for every token in its vocabulary, based on the text so far. The full walkthrough is in how LLMs work.
  2. Temperature scaling: Each logit is divided by the temperature value, which widens or compresses the differences between competing scores.
  3. Softmax: The scaled logits are converted into probabilities by the softmax function, which makes every value positive and the total equal to 1.
  4. Top-p filtering: Candidates are sorted by probability, and the smallest leading group whose cumulative total reaches the top-p value is retained.
  5. Sampling: One token is drawn randomly from the retained candidates, weighted by their renormalised probabilities.
  6. Repetition: The selected token is appended to the input sequence, and the process repeats for the following position until the response is complete.

Temperature Scaling and Nucleus Sampling

Temperature is a single division inside the softmax function, and top-p is a selection rule that keeps the smallest group of probable tokens. Both operations act on the same list of logits produced by the model.

Softmax Temperature Formula

  • : the logit, meaning the unnormalised score of token .
  • : the temperature, a positive number such as 0.5, 1 or 2.
  • : the probability of token after temperature scaling.

The effect is easiest to understand through the ratio between two tokens, which depends only on the difference between their logits divided by the temperature.

A low temperature enlarges every difference, so the leading token dominates the distribution. A high temperature compresses every difference, so the probabilities move towards equal values. As the temperature approaches zero, sampling becomes equivalent to greedy decoding.

The Top-p (Nucleus) Set

  • : the top-p threshold, commonly a value such as 0.9.
  • : the probability of a token after temperature scaling.
  • : the nucleus, constructed by adding tokens in descending order of probability until their cumulative total reaches the threshold.
  • : the renormalised probability that is used for the final random selection.

Worked Example at T = 0.5, 1 and 2

The logits of range( and enumerate( are 4.0 and 3.2, so their difference is 0.8 before temperature scaling.

  1. Low temperature: At 0.5, range( becomes almost five times as probable as enumerate(, so completions become highly consistent.
  2. Neutral temperature: At 1, the original distribution is preserved, and the leading token is a little more than twice as probable.
  3. High temperature: At 2, the ratio falls below one and a half, so alternative completions appear much more frequently.
  4. Nucleus construction: At temperature 1, the running total first reaches the threshold after three tokens, so the nucleus contains three candidates.

The program below calculates all five probabilities at each temperature and the nucleus size for a threshold of 0.9.

Python
import math

# Logits for the token after "for i in " (illustrative, as above)
logits = {"range(": 4.0, "enumerate(": 3.2, "items": 2.5, "zip(": 1.4, "reversed(": 0.6}

def softmax_t(scores, t):
    exps = {tok: math.exp(z / t) for tok, z in scores.items()}
    total = sum(exps.values())
    return {tok: e / total for tok, e in exps.items()}

def nucleus(probs, p):
    # Smallest set of top tokens whose probabilities add up to at least p
    kept, running = [], 0.0
    for tok, pr in sorted(probs.items(), key=lambda x: -x[1]):
        kept.append(tok)
        running += pr
        if running >= p:
            return kept, running

print("token        T=0.5  T=1.0  T=2.0")
table = {t: softmax_t(logits, t) for t in (0.5, 1.0, 2.0)}
for tok in logits:
    print(f"{tok:11}", "  ".join(f"{table[t][tok]:.3f}" for t in table))
for t, probs in table.items():
    kept, mass = nucleus(probs, 0.9)
    print(f"T={t}: top-p 0.9 keeps {len(kept)} tokens (mass {mass:.3f})")
Output
token        T=0.5  T=1.0  T=2.0
range(      0.795  0.562  0.385
enumerate(  0.160  0.252  0.258
items       0.040  0.125  0.182
zip(        0.004  0.042  0.105
reversed(   0.001  0.019  0.070
T=0.5: top-p 0.9 keeps 2 tokens (mass 0.955)
T=1.0: top-p 0.9 keeps 3 tokens (mass 0.940)
T=2.0: top-p 0.9 keeps 4 tokens (mass 0.930)
Softmax temperature on five code completion tokens: a low temperature sharpens the probabilities and a high temperature flattens themGrouped bars show the probability of each candidate token after the Python code for i in, at temperatures 0.5, 1 and 2. range( falls from 0.795 to 0.562 to 0.385 as the temperature rises, while reversed( rises from 0.001 to 0.019 to 0.070. enumerate( is 0.160, 0.252 and 0.258, items is 0.040, 0.125 and 0.182, and zip( is 0.004, 0.042 and 0.105. The logits are illustrative, not from a real model.T = 0.5T = 1T = 20.40.80probabilityrange(enumerate(itemszip(reversed(
Softmax temperature on five code completion tokens: a low temperature sharpens the probabilities and a high temperature flattens them
  • Ratios confirm the formula: At temperature 0.5, the printed probabilities give a ratio of approximately 4.97, which matches the exponential calculation within rounding.
  • Nucleus size follows temperature: The same threshold keeps two tokens at temperature 0.5 but four tokens at temperature 2, because a flatter distribution requires more candidates to reach the same cumulative probability.

Because temperature also changes how many tokens top-p retains, adjusting one parameter at a time makes its effect on code completions considerably easier to measure.

Example: Temperature Scaling and Top-p Filtering in Python

The program below applies both settings to five illustrative scores for the token after for i in .

Python
import math

# Raw scores (logits) for the token after "for i in " (illustrative)
logits = {"range(": 4.0, "enumerate(": 3.2, "items": 2.5, "zip(": 1.4, "reversed(": 0.6}

def softmax(scores, temperature):
    scaled = {t: s / temperature for t, s in scores.items()}
    top = max(scaled.values())
    exps = {t: math.exp(s - top) for t, s in scaled.items()}
    total = sum(exps.values())
    return {t: e / total for t, e in exps.items()}

def top_p(probs, p):
    # Keep the smallest set of top tokens whose probabilities add up to p
    kept, running = {}, 0.0
    for t, pr in sorted(probs.items(), key=lambda x: -x[1]):
        kept[t] = pr
        running += pr
        if running >= p:
            break
    total = sum(kept.values())
    return {t: pr / total for t, pr in kept.items()}

def show(label, probs):
    print(label, ", ".join(f"{t} {pr:.2f}" for t, pr in probs.items()))

for temp in (0.5, 1.0, 1.5):
    show(f"T={temp}:", softmax(logits, temp))
show("T=1.0, top_p=0.9:", top_p(softmax(logits, 1.0), 0.9))
Output
T=0.5: range( 0.79, enumerate( 0.16, items 0.04, zip( 0.00, reversed( 0.00
T=1.0: range( 0.56, enumerate( 0.25, items 0.13, zip( 0.04, reversed( 0.02
T=1.5: range( 0.45, enumerate( 0.26, items 0.16, zip( 0.08, reversed( 0.05
T=1.0, top_p=0.9: range( 0.60, enumerate( 0.27, items 0.13
  • Low temperature: At 0.5, range( rises to 0.79 and the two weakest tokens fall to almost zero, so completions become consistent.
  • High temperature: At 1.5, the weakest token reversed( rises from 0.02 to 0.05, so unusual completions appear more often.
  • Top-p cut: With top-p 0.9, the first two tokens reach only 0.81, so items is also kept, while zip( and reversed( are removed.

Applications of Temperature and Top-p

  • Code completion: Low temperature keeps inline suggestions stable and consistent with conventional coding patterns.
  • Structured output: Low values reduce formatting errors when a model must return structured JSON output.
  • Test data generation: Higher values produce more diverse sample inputs, identifiers and edge cases.
  • Brainstorming: Higher values generate a wider variety of variable names, commit messages or architectural alternatives.
  • Multiple candidates: Systems that generate several solutions and select one require some randomness to produce distinct candidates.
  • Evaluation runs: Fixed low values make repeated evaluation runs easier to compare, as explained in AI agent evaluation.

Advantages

  • Simple configuration: Two numeric parameters change output behaviour without fine-tuning or retraining the model.
  • Per-request tuning: Each API call can use different values for different tasks within the same application.
  • Fewer anomalous tokens: Top-p removes the long tail of improbable tokens that produce irrelevant or incoherent output.
  • Predictability on demand: Low temperature supports repeatable responses in automated pipelines and continuous integration checks.

Limitations

  • No factual guarantee: A temperature of 0 still produces confident errors, a problem covered in LLM hallucination.
  • Not fully deterministic: Identical requests can still differ slightly because of hardware and batching effects.
  • Interacting settings: Changing both settings at once makes results hard to reason about, so providers often advise tuning one.
  • Model-specific behaviour: The same value can behave differently across models, so tuned settings rarely transfer between them.
  • Restricted on some models: Some reasoning models ignore these parameters or fix them at predetermined values.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does a temperature below 1.0 do to next-token probabilities?

Frequently Asked Questions

What is a good temperature for code generation?

Code generation usually works well with a low temperature, because correct code follows common patterns. Higher values are useful only when several different candidate solutions are needed.

Temperature vs top-p: what is the difference?

Temperature changes the shape of the whole probability distribution before sampling. Top-p removes the unlikely tokens and samples only from the most likely group that covers a set share of probability.

Does temperature 0 make an LLM deterministic?

Temperature 0 makes the model pick the most likely token at each step, so outputs become very consistent. Small differences can still appear because of how requests are batched and computed on hardware.

Is top-p the same as nucleus sampling?

Yes, top-p sampling is also called nucleus sampling. The nucleus is the smallest group of top tokens whose probabilities add up to the chosen value.

Does a higher temperature make an LLM more creative?

A higher temperature makes less likely tokens more common, so output becomes more varied. It does not add knowledge or reasoning, and very high values often produce incoherent text.