Context Window in LLM
The context window is the maximum number of tokens a large language model can process in one request, counting the instructions, attached documents, conversation history and generated output together. Text beyond this limit is not seen by the model, so applications must select, truncate or summarise their input to fit.
- Token-based limit: The window is measured in tokens rather than characters or words, so its practical capacity depends on tokenization in LLMs.
- Shared budget: Input and generated output draw on the same allocation, so an extensive prompt leaves fewer tokens available for the response.
- Stateless requests: The model retains nothing between API calls, so a chat application resends the relevant history with every request.
- Model-dependent size: LLM context length, another name for the window size, ranges from a few thousand tokens to over one million tokens in current models.
- Working memory: The window is the only short-term memory available to the model during inference.
For example, a developer assistant asked to diagnose a failing service may receive a 5,000-line log, but only the most recent lines that fit the context window, including the stack trace, can be sent.
Key Characteristics of the Context Window
- Hard limit: Requests that exceed the window are rejected by the API or silently truncated by the application.
- Attention cost: Standard attention compares every token with every other token, so computation grows faster than the input length.
- Position effects: Models often recall information at the start and end of a long input more reliably than information in the middle.
- Latency and price: Longer inputs increase the time to the first generated token and the cost of each request.
- Separate output cap: Many models also enforce a smaller, independent limit on the number of generated tokens.
- Memory footprint: Serving long contexts requires a key-value cache whose size grows with every token held in the window.
How the Context Window Works
- Assembly: The application combines system instructions, conversation history, retrieved documents and the new user message into one prompt.
- Counting: A tokenizer calculates the token count of the assembled prompt before transmission.
- Reservation: A portion of the window is reserved for the expected length of the response.
- Fitting: If the prompt exceeds the remaining budget, the application drops, truncates or summarises the least important content, a practice known as context engineering.
- Processing: The model applies attention across every token inside the window, as described in transformer architecture.
- Generation: Generated tokens are appended to the window until the response completes or the token limit is reached.
Attention Cost and KV Cache Memory Formulas
Two costs grow with the number of tokens in the window: attention computation grows with the square of that number, and the memory of the key-value (KV) cache grows in direct proportion to it.
Attention Cost Grows with n Squared
- : the number of tokens currently inside the context window.
- : the query and key matrices, each containing one row per token.
- : the dimension of each query and key vector.
- : the score matrix, containing one row and one column for every token.
In words, every token is compared with every other token, so doubling the input length quadruples the number of attention scores. A causal mask removes approximately half of these comparisons, but the growth remains quadratic.
KV Cache Size Formula
During generation, the model stores the key and value vectors of every previous token, so it avoids recalculating them for each newly generated token.
- : the factor for storing one key vector and one value vector.
- : the number of transformer layers in the model.
- : the number of heads that store keys and values, which equals every head in standard attention and fewer in multi-query or grouped-query attention.
- : the dimension of each attention head.
- : the number of tokens currently held in the cache during generation.
- : the number of bytes for each stored value, which is 2 for 16-bit precision.
The cache therefore grows linearly, because every additional token contributes an identical, predetermined amount of accelerator memory to each individual request.
Worked Example with Round Model Sizes
The sizes below are round, illustrative numbers chosen for convenient arithmetic, rather than the published specification of any particular model.
- Per-token memory: Every cached token requires half a mebibyte of memory across all layers and attention heads of this illustrative configuration.
- Stack trace and source file: A prompt containing a long traceback and its related module needs almost four gibibytes of cache.
- Repository-scale prompt: Sixteen times as many tokens require sixteen times the memory, while the score matrix grows by a factor of 256.
- Fewer key-value heads: Sharing keys and values across heads reduces the cache in proportion to the number of heads removed.
# Round, illustrative model sizes (not a specific real model)
LAYERS = 32
KV_HEADS = 32 # heads that store keys and values
HEAD_DIM = 128
BYTES = 2 # 16-bit numbers
def kv_cache_bytes(tokens, kv_heads=KV_HEADS):
# 2 (keys and values) x layers x heads x head_dim x tokens x bytes
return 2 * LAYERS * kv_heads * HEAD_DIM * tokens * BYTES
GIB = 1024 ** 3
print(f"Per token: {kv_cache_bytes(1) // 1024} KiB")
print("tokens KV cache scores per head and layer")
for n in (1_000, 8_000, 32_000, 128_000):
print(f"{n:7,} {kv_cache_bytes(n) / GIB:6.2f} GiB {n * n:>17,}")
# Fewer key-value heads shrink the cache by the same factor
n = 32_000
for heads in (32, 8, 1):
print(f"KV heads = {heads:2}, {n:,} tokens: {kv_cache_bytes(n, heads) / GIB:6.2f} GiB")Per token: 512 KiB
tokens KV cache scores per head and layer
1,000 0.49 GiB 1,000,000
8,000 3.91 GiB 64,000,000
32,000 15.62 GiB 1,024,000,000
128,000 62.50 GiB 16,384,000,000
KV heads = 32, 32,000 tokens: 15.62 GiB
KV heads = 8, 32,000 tokens: 3.91 GiB
KV heads = 1, 32,000 tokens: 0.49 GiB- Memory is linear, computation is quadratic: Moving from 32,000 to 128,000 tokens multiplies the cache by four and the attention scores by sixteen.
- Head sharing is the principal lever: Reducing the number of key-value heads shrinks the cache without shortening the context length.
These formulas explain why long prompts increase latency and price: computation grows with the square of the input, and cache memory limits how many long requests one accelerator can serve concurrently.
Example: Fitting a Log into the Context Window
The program below keeps the newest lines of a Python error log that fit a 60-token window after reserving 15 tokens for the answer.
import re
def count_tokens(text):
# Rough stand-in for a real tokenizer: words, numbers and symbols
return len(re.findall(r"\w+|[^\w\s]", text))
WINDOW = 60 # total tokens the model accepts
RESERVED_OUTPUT = 15 # tokens set aside for the answer
instruction = "Explain the root cause of this Python traceback."
log = [f"INFO worker {n} processed batch {n}" for n in range(1, 9)] + [
"Traceback (most recent call last):",
'File "app.py", line 42, in load',
"KeyError: 'user_id'",
]
def fit_to_window(instruction, lines):
budget = WINDOW - RESERVED_OUTPUT - count_tokens(instruction)
kept = []
# Walk backwards so the newest lines, including the error, are kept
for line in reversed(lines):
cost = count_tokens(line)
if cost > budget:
break
kept.insert(0, line)
budget -= cost
return kept, budget
total = count_tokens(instruction) + sum(count_tokens(l) for l in log)
kept, left = fit_to_window(instruction, log)
print(f"Full prompt: {total} tokens, window: {WINDOW}, reserved for output: {RESERVED_OUTPUT}")
print(f"Kept {len(kept)} of {len(log)} log lines, {left} tokens unused:")
for line in kept:
print(" ", line)Full prompt: 82 tokens, window: 60, reserved for output: 15
Kept 4 of 11 log lines, 5 tokens unused:
INFO worker 8 processed batch 8
Traceback (most recent call last):
File "app.py", line 42, in load
KeyError: 'user_id'- Newest lines survive: Walking backwards preserves the traceback and the
KeyError, which are the lines a model needs to diagnose the failure. - Output is budgeted first: Reserving 15 tokens prevents the prompt from consuming the space required for the explanation.
- Production systems do more: Real assistants count with the model's own tokenizer and often summarise older lines instead of discarding them.
Applications of a Large Context Window
- Whole-file code review: Reviewing an entire module or a complete pull request diff in a single request.
- Log analysis: Examining lengthy stack traces and surrounding log entries during incident investigation.
- Multi-turn assistants: Retaining earlier conversation turns so that follow-up questions are interpreted correctly.
- Document question answering: Supplying API references directly, or selecting relevant passages through RAG.
- Agent workflows: Accumulating tool results and intermediate reasoning steps for an AI agent.
- Few-shot prompting: Including several worked examples that demonstrate the required output structure.
Advantages
- Fewer retrieval steps: A large window can accommodate entire documents without a separate retrieval pipeline.
- Better coherence: Longer history lets the model remain consistent across an extended conversation or a multi-file codebase.
- Richer instructions: Detailed specifications, coding conventions and examples fit alongside the primary task.
- Simpler engineering: Moderately sized inputs require less chunking, ranking and summarisation logic.
Limitations
- Higher cost: Every token in the window is processed and billed on every request, including repeated conversation history.
- Increased latency: Lengthy prompts increase the latency before the first generated token appears.
- Degraded recall: Relevant details positioned in the middle of very long inputs are more likely to be overlooked.
- Distracting content: Loosely related text can mislead the model and increase LLM hallucination.
- No persistence: Content leaves the model's view once it is dropped from the window, so long-term memory requires external storage.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In what unit is a context window measured?
Frequently Asked Questions
What happens when a prompt exceeds the context window?
The API usually rejects the request with an error, or the application truncates the input before sending it. Any content that is removed is invisible to the model, so it cannot influence the response.
Does the context window include the model's output?
Yes. Input and generated output share the same token limit, so a very long prompt reduces the space available for the answer. Many applications reserve a fixed number of tokens for output.
Is a bigger context window always better?
Not always. Longer inputs cost more, respond more slowly and can bury important details, so a smaller set of relevant content often produces better answers.
How is the context window different from memory?
The context window holds only what is sent in the current request. Persistent memory requires the application to store information externally and add the relevant parts to later prompts.
Related Articles
- Tokens and Tokenization in LLMLearn how tokenization in LLMs splits text and code into subword tokens and why token counts drive cost and context, with a Python tokenizer example.
- Transformer Architecture ExplainedTransformer architecture explained: self-attention, query, key and value vectors, multi-head attention and stacked layers, with Python attention code.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- What is Context EngineeringLearn what context engineering is: selecting, ranking and fitting information into an LLM context window, with a Python token budget example and limits.
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.