…
Skip to content
Topics
On this page

Transformer Architecture Explained

Transformer architecture is a neural network design that processes a whole sequence of tokens in parallel and uses an attention mechanism called self-attention to relate every token to the others. Introduced in the 2017 paper "Attention Is All You Need", it is the foundation of almost every modern large language model.

  • Self-attention: Each token computes relevance weights over the other tokens and builds a new representation from their weighted combination.
  • Parallel processing: All positions in a sequence are processed simultaneously during training, unlike recurrent networks that read one token at a time.
  • Stacked layers: A model repeats the same transformer block dozens of times, and each layer refines the representations produced by the previous one.
  • Positional information: Position encodings add word-order information, because attention by itself treats the input as an unordered set.
  • Decoder-only variant: Most generative LLMs use only the decoder half with a causal mask, so each token attends exclusively to earlier tokens.
Inside one transformer blockToken vectors with position information enter a transformer block. Inside it, masked self-attention builds query, key and value vectors, then an add and normalise step, a feed-forward network and a second add and normalise step follow. Dashed residual connections carry each sublayer's input around it. The block is repeated N times, and the final output becomes next-token scores. The layout is a simplified sketch.Tokens +positionsTransformer block, repeated N timesMaskedattentionAdd andnormFeed-forwardAdd andnormQ, K, VDashed arcs: residual connectionsNext-tokenscores
Inside one transformer block

For example, when a code assistant reads total = total + item.price, attention lets the token price draw much of its context from item, which identifies the object that owns the attribute.

Key Characteristics of Transformer Architecture

  • Query, key and value vectors: Each token is projected into three vectors: the query describes what it seeks, the key describes what it offers and the value carries its content.
  • Scaled dot-product attention: Attention scores are dot products of queries and keys divided by the square root of the vector dimension, which keeps the softmax inputs in a stable range.
  • Multi-head attention: Several attention heads run in parallel, so one head can track syntax while another tracks variable references.
  • Feed-forward sublayer: After attention, each token passes through a small fully connected network that transforms its representation independently.
  • Residual connections and normalisation: Each sublayer adds its output to its input and normalises the result, which stabilises the training of very deep stacks.
  • Quadratic attention cost: Every token attends to every permitted token, so computation grows with the square of the sequence length.

How Transformer Architecture Works

  1. Token embedding: Input token IDs are mapped to learned vectors, as described in embeddings in LLMs.
  2. Position encoding: Positional information is added to or combined with each vector, so the model can distinguish a - b from b - a.
  3. Projection: Each layer multiplies every token vector by three learned matrices to produce its query, key and value vectors.
  4. Attention weighting: Each query is compared with the keys of permitted tokens, and softmax converts the scaled scores into weights that sum to 1.
  5. Value mixing: The new representation of each token is the weighted sum of the value vectors, followed by the feed-forward sublayer.
  6. Repetition across layers: The block repeats many times, and the output of the final layer is converted into next-token probabilities, the last stage of how LLMs work.

Scaled Dot-Product Attention, Step by Step

Attention reduces to one compact formula: compare every query with every key, convert the resulting scores into weights with softmax, and use those weights to combine the value vectors.

The Scaled Dot-Product Attention Formula

  • : the query matrix, containing one row per token that describes the information the token is searching for.
  • : the key matrix, containing one row per token that describes the information the token can provide.
  • : the value matrix, containing one row per token with the content that is eventually combined.
  • : the dimension of each query and key vector.
  • : the matrix of dot products, in which each entry measures how closely one query matches one key.
  • : the causal mask, which is zero at permitted positions and negative infinity at every later position.

In words, this attention formula calculates all query and key similarities simultaneously, scales them, normalises each row with softmax, and multiplies the resulting weights by the value vectors. Encoder models apply it without the mask, while decoder-only models add the mask so that no token can observe future positions.

Why Divide by the Square Root of the Key Dimension

When the components of a query and a key are independent, with mean zero and variance one, the variance of their dot product equals the vector dimension.

Typical scores therefore grow with the square root of the dimension. Large scores push softmax towards a single weight near one, where the gradients become extremely small and training slows considerably. Dividing by the square root of the dimension restores unit variance for every head size.

Multi-Head Attention

  • : the input matrix, containing one embedding row for every token in the sequence.
  • : the number of attention heads, each with its own learned projection matrices.
  • : the output projection, which combines the concatenated head outputs into the model dimension.

Each head operates in a smaller subspace, so several heads together cost approximately the same as one full-width head, yet each head can specialise in a different relationship, such as syntax or variable references.

Worked Example: Attention over item.price

The code item.price produces three tokens with hand-made two-dimensional vectors, so the scaling factor is the square root of two. The calculation below follows the row for price, whose query vector is .

  1. Similarity scores: The query of price aligns most closely with the key of item, which identifies the object that owns the attribute.
  2. Normalised weights: Softmax assigns most of the attention to item and very little to the dot operator.
  3. Weighted combination: The output vector is dominated by the value of item, which is exactly the contextual information the token requires.
  4. Causal masking: The first token, item, can observe only itself, so its complete weight remains on its own position.
Python
import math

tokens = ["item", ".", "price"]
# Tiny 2-dimensional query, key and value vectors (hand-made, d_k = 2)
Q = [[1.0, 0.0], [0.0, 1.0], [2.0, 1.0]]
K = [[2.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
V = [[1.0, 0.0], [0.0, 1.0], [0.5, 0.5]]
d_k = 2

def softmax(xs):
    top = max(xs)
    exps = [math.exp(x - top) for x in xs]
    return [e / sum(exps) for e in exps]

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

# Causal mask: row i may only look at columns 0..i
print("Attention weights (rows attend to columns):")
weights = []
for i, q in enumerate(Q):
    scores = [dot(q, K[j]) / math.sqrt(d_k) for j in range(i + 1)]
    w = softmax(scores) + [0.0] * (len(tokens) - i - 1)
    weights.append(w)
    print(f"  {tokens[i]:5}", " ".join(f"{x:.2f}" for x in w))

# Output of each token: weighted sum of the value vectors
for i, w in enumerate(weights):
    out = [sum(w[j] * V[j][c] for j in range(len(V))) for c in range(2)]
    print(f"Output for {tokens[i]!r}: [{out[0]:.2f}, {out[1]:.2f}]")

# Why divide by sqrt(d_k): unscaled scores grow with d_k and saturate softmax
row = [dot(Q[2], k) for k in K]
print("price, scaled by sqrt(2):", [round(x, 2) for x in softmax([s / math.sqrt(2) for s in row])])
print("price, scores 8x larger: ", [round(x, 2) for x in softmax([s * 8 / math.sqrt(2) for s in row])])
Output
Attention weights (rows attend to columns):
  item  1.00 0.00 0.00
  .     0.33 0.67 0.00
  price 0.62 0.07 0.31
Output for 'item': [1.00, 0.00]
Output for '.': [0.33, 0.67]
Output for 'price': [0.77, 0.23]
price, scaled by sqrt(2): [0.62, 0.07, 0.31]
price, scores 8x larger:  [1.0, 0.0, 0.0]
Causal attention weights for the code item.price: each row is a query token and each column is a key tokenA three by three heatmap of scaled dot-product attention weights with a causal mask, using hand-made 2-dimensional vectors. Row item: 1.00 on item, later columns masked. Row dot: 0.33 on item and 0.67 on dot, last column masked. Row price: 0.62 on item, 0.07 on dot and 0.31 on price. Each row sums to 1, darker cells carry more weight, and hatched cells are masked future positions. The vectors and weights are illustrative.item.priceitem.price1.000.330.670.620.070.31
Causal attention weights for the code item.price: each row is a query token and each column is a key token
  • Rows are probability distributions: Every row of weights sums to one, and masked positions receive exactly zero.
  • Scaling keeps weights balanced: When identical scores are multiplied by eight, as unscaled scores typically are for a dimension of 64, softmax places the entire weight on one token.

Head size, head count and the quadratic score matrix determine the computation and memory of every layer, which is why long inputs strain the context window and increase the cost of every request.

Example: Computing Attention Weights in Python

The program below computes scaled dot-product attention weights for the token price in one line of Python code, using hand-made query and key vectors.

Python
import math

# Tokens of: total = total + item.price
tokens = ["total", "=", "total", "+", "item", ".", "price"]
# Hand-made 4-dimensional query and key vectors (a trained model learns these)
keys = {
    "total": [1.0, 0.2, 0.0, 0.3],
    "=": [0.0, 1.0, 0.0, 0.0],
    "+": [0.1, 0.9, 0.1, 0.0],
    "item": [0.3, 0.0, 1.0, 0.8],
    ".": [0.0, 0.6, 0.2, 0.0],
    "price": [0.4, 0.0, 0.7, 1.0],
}
# The query of the last token, "price", looks for the object that owns it
query = [0.2, 0.0, 1.2, 0.9]

def softmax(xs):
    exps = [math.exp(x) for x in xs]
    return [e / sum(exps) for e in exps]

d = len(query)
# Scaled dot-product attention: score = (q . k) / sqrt(d)
scores = [sum(q * k for q, k in zip(query, keys[t])) / math.sqrt(d) for t in tokens]
weights = softmax(scores)

for t, s, w in zip(tokens, scores, weights):
    print(f"{t:6} score={s:5.2f}  weight={w:.2f}  {'#' * round(w * 40)}")
print("Weights sum to", round(sum(weights), 2))
Output
total  score= 0.24  weight=0.12  #####
=      score= 0.00  weight=0.09  ####
total  score= 0.24  weight=0.12  #####
+      score= 0.07  weight=0.10  ####
item   score= 0.99  weight=0.25  ##########
.      score= 0.12  weight=0.10  ####
price  score= 0.91  weight=0.23  #########
Weights sum to 1.0
  • Relevant tokens receive more weight: item has the highest score and takes 25 percent of the attention, because its key vector aligns most closely with the query from price.
  • Weights form a distribution: Softmax makes every weight positive and the total equal to 1, so the result is a weighted average of value vectors.
  • Learned in practice: A trained transformer learns the query and key projections from data, and each of its many heads produces a separate set of weights.

Applications of Transformer Architecture

  • Large language models: The transformer in LLMs generates natural language and source code for conversational assistants and developer tools.
  • Code assistants: Completing functions, explaining stack traces and summarising pull requests across long source files.
  • Machine translation: The original task for which the encoder-decoder transformer was designed.
  • Embedding models: Encoder transformers produce vectors for semantic search and RAG.
  • Vision transformers: Image patches are treated as tokens for classification, detection and visual understanding.
  • Speech recognition: Audio frames are encoded as token sequences and transcribed into written text.

Advantages

  • Parallel training: Entire sequences are processed simultaneously, which suits GPUs and other specialised accelerators.
  • Long-range dependencies: Any token can attend directly to any earlier token, regardless of distance.
  • Scalability: Performance has improved consistently as parameters, training data and compute have increased.
  • Modality independence: The same block processes text, code, images and audio once they are converted into tokens.

Limitations

  • Quadratic cost: Attention compute and memory grow with the square of the input length, which constrains the context window.
  • Inference memory: The key-value cache for long prompts consumes substantial accelerator memory.
  • No inherent ordering: Sequence order must be supplied separately through positional encodings or rotary position embeddings.
  • Data requirements: Transformers typically require very large training corpora and considerable compute to reach competitive accuracy.
  • Limited interpretability: Attention weights show where a model looks but do not fully explain its output.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. Why are attention scores divided by the square root of the vector dimension?

Frequently Asked Questions

What is the transformer architecture in simple terms?

A transformer model is a neural network that reads all tokens of an input at once and uses attention to decide how much each token should influence the others. Stacking many such layers lets the model build a detailed representation of the text.

What is self-attention?

Self-attention is the operation in which every token in a sequence compares its query vector with the key vectors of other tokens. The resulting weights decide how their value vectors are combined into a new representation.

Why did transformers replace recurrent neural networks?

Recurrent networks process tokens one after another, which slows training and weakens long-range connections. Transformers process whole sequences in parallel and let any token attend directly to any earlier token.

What is the difference between an encoder and a decoder transformer?

An encoder lets each token attend to the entire input and is suited to understanding tasks such as classification and embeddings. A decoder uses a causal mask so each token sees only earlier tokens, which suits text generation.

Is GPT based on the transformer architecture?

Yes. GPT stands for Generative Pre-trained Transformer, and it uses a decoder-only transformer trained to predict the next token.