…
Skip to content
Topics
On this page

How LLMs Work

How LLMs work can be described in two phases: training, in which a model adjusts its parameters by predicting the next token across large text and code corpora, and inference, in which the trained model generates output for a prompt. Both phases share one pipeline of tokenization, embeddings, attention layers and a final scoring step.

  • Training phase: The model processes trillions of tokens and adjusts billions of parameters to reduce its prediction error.
  • Inference phase: The trained model runs a forward pass, one computation through all of its layers, for every token it generates.
  • Shared pipeline: Text passes through a tokenizer, an embedding layer, a stack of transformer layers and an output layer in both phases.
  • Autoregressive generation: Each generated token is appended to the input before the next forward pass, so output is produced strictly from left to right.
  • Frozen weights at inference: Parameters do not change while the model answers, so new information must arrive through the prompt itself.
How LLMs work: training and inferenceTwo rows. The training row runs from a code corpus to predicting the next token, measuring the loss and updating the weights. The trained weights flow down into the inference row, which runs from tokens to embeddings, attention layers, next-token probabilities and decoding. A dashed arrow returns from decoding to the tokens, because each generated token is appended and the pipeline runs again. The diagram is a simplified sketch, not a real model.TrainingCode corpusPredict next tokenLossUpdate weightsTrained weightsInferenceTokensEmbeddingsAttention layersProbabilitiesDecode
How LLMs work: training and inference

For example, when a developer assistant explains a Python stack trace, it tokenizes the trace, passes the tokens through dozens of transformer layers and decodes the explanation one token at a time.

Key Characteristics of the LLM Pipeline

  • Self-supervised objective: Training labels come from the text itself, because the next token in every sequence is the target the model must predict.
  • Loss minimisation: A loss function called cross-entropy measures prediction error, and gradient descent adjusts the parameters to lower it.
  • Staged training: Pre-training builds general language ability, and later stages such as instruction tuning and preference tuning shape the model into a cooperative assistant.
  • Layered computation: Dozens of stacked layers of the transformer architecture progressively refine the representation of each token.
  • Parallel training, sequential inference: Training processes every position of a sequence simultaneously, while inference must generate tokens one after another.
  • Cached computation: A key-value cache stores intermediate attention results for earlier tokens, so each new token avoids recomputing the entire prompt.

How LLMs Work Step by Step

The steps below show how large language models work from raw data to generated text. They add the training stages to the inference overview given for a large language model.

  1. Data collection: Engineers assemble web pages, books, documentation and public source code, then filter duplicates and low-quality text.
  2. Tokenization: A tokenizer converts the text into integer IDs from a fixed vocabulary, as described in tokenization in LLMs.
  3. Embedding: An embedding layer maps each ID to a learned vector, the numerical representation covered in embeddings in LLMs.
  4. Attention layers: Each layer lets every token gather information from earlier tokens, then applies a feed-forward network to transform the combined result.
  5. Output scoring: The final layer converts the vector of the last token into a score for every vocabulary entry, and a softmax function turns those scores into probabilities.
  6. Training update: During training, the loss compares these probabilities with the true next token, and backpropagation adjusts every parameter by a small amount.
  7. Decoding: During inference, a decoding strategy such as greedy selection or sampling with temperature and top-p picks one token, which is appended before the next forward pass.

Formulation of Next-Token Prediction

Two formulas drive the whole pipeline: the softmax function converts the model's raw output scores into next-token probabilities, and the cross-entropy loss measures how much probability the model assigned to the correct token during training.

Softmax: From Logits to Probabilities

The softmax function exponentiates every logit, which makes each value positive, and then divides each result by the total so that the probabilities add up to exactly 1.

  • : the logit, meaning the unnormalised score that the output layer assigns to token .
  • : the number of distinct tokens in the tokenizer vocabulary.
  • : the resulting probability of token , which is always positive and never larger than 1.

A larger logit always produces a larger probability, and a difference of one unit between two logits makes one token about 2.7 times as probable as the other.

Cross-Entropy Loss in LLM Training

The cross-entropy loss for a single position is the negative natural logarithm of the probability that the model assigned to the true next token. For a complete training sequence, the objective is the average of these individual losses across every predicted position.

  • : the true next token, taken directly from the training text.
  • : the number of predicted positions in the training sequence.
  • : the conditional probability of token , given every preceding token.

The loss equals zero only when the correct token receives the entire probability, and it increases steeply as that probability approaches zero.

Worked Example: One Python Token

A code model completes def add(a, b): return a and produces logits for four candidate operators and punctuation marks.

  1. Exponentiation: Each logit becomes a positive value, and the four exponentials add up to a common denominator.
  2. Normalisation: Dividing by that denominator gives the addition operator about 61 percent of the probability.
  3. Low loss for a likely token: When the training text actually continues with the addition operator, the penalty is small.
  4. High loss for an unlikely token: A training text that continued with a comma would produce a penalty about seven times larger.

The program below reproduces these calculations and then averages the loss across three consecutive positions.

Python
import math

# Logits for the token after "return a " in: def add(a, b): return a
logits = {"+": 2.0, "-": 1.0, "*": 0.5, ",": -1.0}

# Softmax: exponentiate each logit, then divide by the total
exps = {t: math.exp(z) for t, z in logits.items()}
total = sum(exps.values())
probs = {t: e / total for t, e in exps.items()}

print(f"Sum of exponentials: {total:.3f}")
for t, p in probs.items():
    print(f"  {t!r}: logit {logits[t]:4.1f}  exp {exps[t]:.3f}  p {p:.3f}")

# Cross-entropy loss: minus the log of the probability of the true token
for target in ["+", ","]:
    loss = -math.log(probs[target])
    print(f"True token {target!r}: loss = -ln({probs[target]:.3f}) = {loss:.3f}")

# Loss over a sequence: the mean of the per-token losses
seq_probs = [0.60, 0.90, 0.25]
mean_loss = -sum(math.log(p) for p in seq_probs) / len(seq_probs)
print(f"Mean loss over 3 tokens: {mean_loss:.3f}")
Output
Sum of exponentials: 12.124
  '+': logit  2.0  exp 7.389  p 0.609
  '-': logit  1.0  exp 2.718  p 0.224
  '*': logit  0.5  exp 1.649  p 0.136
  ',': logit -1.0  exp 0.368  p 0.030
True token '+': loss = -ln(0.609) = 0.495
True token ',': loss = -ln(0.030) = 3.495
Mean loss over 3 tokens: 0.667
Softmax probabilities for the token after return a, and the cross-entropy loss curve L = -ln pLeft: a bar chart of the softmax probabilities of four candidate tokens after the Python code return a. The plus sign, the true token, has probability 0.61, the minus sign 0.22, the asterisk 0.14 and the comma 0.03. Right: the curve of the loss, minus the natural log of the probability of the true token, falling from about 3.9 at probability 0.02 to 0 at probability 1. Two points are marked: probability 0.61 gives loss 0.50, and probability 0.03 gives loss 3.50. The logits and probabilities are illustrative, not from a real model.Softmax probabilities+ 0.61- 0.22* 0.14, 0.03true tokenLoss L = -ln pp 0.03, loss 3.50p 0.61, loss 0.50p of the true token (0 to 1)
Softmax probabilities for the token after return a, and the cross-entropy loss curve L = -ln p
  • Confident, correct predictions are inexpensive: A probability of 0.90 contributes only 0.105, whereas 0.25 contributes 1.386, so inaccurate predictions dominate the average.
  • Direction of training: Backpropagation adjusts the weights so that the logit of the correct token increases and the competing logits decrease.

Gradient descent minimises this average loss across the entire training corpus, so the cross-entropy curve is the primary measurement engineers monitor to decide whether pre-training is still improving a model.

Example: Tracing One Inference Step in Python

The program below traces one forward pass of a toy model with a four-token vocabulary and three-dimensional embeddings.

Python
import math

# A toy model: 4-token vocabulary, 3-dimensional embeddings
vocab = ["return", "total", "names", "0"]
embed = {
    "return": [1.0, 0.0, 0.5],
    "total": [0.2, 1.0, 0.1],
    "names": [0.1, 0.8, 0.9],
    "0": [0.9, 0.1, 0.0],
}

def softmax(xs):
    exps = [math.exp(x) for x in xs]
    return [e / sum(exps) for e in exps]

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

# 1. Tokenization: split the prompt and map each token to an ID
prompt = "return total"
tokens = prompt.split()
print("1. Token IDs:", [vocab.index(t) for t in tokens])

# 2. Embedding: look up one vector per token
vectors = [embed[t] for t in tokens]

# 3. Attention: the last token weighs every token, then mixes their vectors
weights = softmax([dot(vectors[-1], v) for v in vectors])
print("2. Attention weights:", [round(w, 2) for w in weights])
context = [sum(w * v[i] for w, v in zip(weights, vectors)) for i in range(3)]

# 4. Scoring: compare the context vector with every vocabulary embedding
probs = softmax([dot(context, embed[t]) for t in vocab])
print("3. Next-token probabilities:")
for t, p in sorted(zip(vocab, probs), key=lambda x: -x[1]):
    print(f"   {t:7} {p:.2f}")

# 5. Decoding: greedy decoding picks the highest probability
print("4. Greedy pick:", vocab[probs.index(max(probs))])
Output
1. Token IDs: [0, 1]
2. Attention weights: [0.31, 0.69]
3. Next-token probabilities:
   total   0.29
   names   0.28
   return  0.22
   0       0.21
4. Greedy pick: total
  • Every stage is arithmetic: Tokenization produces IDs, attention produces weights that sum to 1, and scoring produces a probability distribution over the vocabulary.
  • Untrained weights give weak predictions: The hand-made vectors spread probability almost evenly, so the winning token total is not a sensible continuation. Training would adjust these numbers.
  • Scale is the main difference: A production model performs the same operations with vocabularies of over 100,000 tokens and thousands of dimensions per vector.

Applications of Knowing How LLMs Work

  • Prompt design: Next-token prediction explains why explicit instructions and worked examples improve output quality.
  • Cost estimation: Token-based pricing and response latency follow directly from the per-token forward pass.
  • Debugging completions: Repetitive or erratic output can often be traced to decoding settings instead of the prompt.
  • Choosing adaptation methods: Frozen inference weights explain why new facts require retrieval or fine-tuning.
  • Context management: The context window limit explains why long source files and logs must be truncated.

Advantages

  • Simple objective: Next-token prediction needs no manual labels, so training can use very large unlabelled corpora.
  • Transferable learning: Knowledge acquired during pre-training carries over to coding, summarisation and question answering.
  • Parallel training: Processing whole sequences simultaneously makes efficient use of modern accelerators.
  • Flexible decoding: The same trained weights support deterministic or varied output through a change of decoding strategy.

Limitations

  • Sequential latency: Long output requires one forward pass per token, which increases response time.
  • Frozen knowledge: The weights encode only information available before the training cut-off.
  • High training cost: Pre-training demands large clusters of accelerators running for weeks.
  • Opaque reasoning: Billions of interacting parameters make it difficult to explain why a particular token was chosen.
  • Error propagation: An incorrect early token becomes part of the input and can steer the remaining output, which contributes to LLM hallucination.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What changes during training but stays fixed during inference?

Frequently Asked Questions

LLM training vs inference: what is the difference?

Training is the phase in which the model adjusts its parameters by predicting the next token across a large corpus. LLM inference is the phase in which the finished model generates output for a prompt without changing its parameters.

Do LLMs learn from the prompts they receive?

A deployed model does not update its parameters while it answers, so a prompt influences only the current response. Whether conversations are later used to train a new model version depends on the provider's data policy.

How do LLMs work in simple terms?

An LLM splits text into tokens, passes them through stacked transformer layers and scores every possible next token. It then selects one token, appends it to the input and repeats, and this autoregressive loop is why long responses take longer to generate.

What is a forward pass in an LLM?

An LLM forward pass is one complete computation through every layer of the model, from token embeddings to the final probability distribution. During inference, the model runs one forward pass for each generated token.

What role does attention play in how LLMs work?

Attention lets each token weigh the relevance of every earlier token and combine information from them. It allows the model to connect a variable name to its definition or a pronoun to its noun across long inputs.