…
Skip to content
Topics
On this page

Context Window in LLM

The context window is the maximum number of tokens a large language model can process in one request, counting the instructions, attached documents, conversation history and generated output together. Text beyond this limit is not seen by the model, so applications must select, truncate or summarise their input to fit.

  • Token-based limit: The window is measured in tokens rather than characters or words, so its practical capacity depends on tokenization in LLMs.
  • Shared budget: Input and generated output draw on the same allocation, so an extensive prompt leaves fewer tokens available for the response.
  • Stateless requests: The model retains nothing between API calls, so a chat application resends the relevant history with every request.
  • Model-dependent size: LLM context length, another name for the window size, ranges from a few thousand tokens to over one million tokens in current models.
  • Working memory: The window is the only short-term memory available to the model during inference.
Fitting a long error log into the context windowThe full log has 11 lines. Lines 1 to 7 are dropped because they do not fit. Lines 8 to 11, which include the traceback and the KeyError, move into a 60-token context window alongside the instruction and a reserve of 15 tokens for the answer. The sizes are illustrative and match the small Python example, not a real model.Full log: 11 linesLines 1 to 7 (oldest)Lines 8 to 11 + tracebackDropped: does not fitContext window: 60 tokensInstructionNewest log linesReserved for outputNewest lines are kept, oldest lines are dropped
Fitting a long error log into the context window

For example, a developer assistant asked to diagnose a failing service may receive a 5,000-line log, but only the most recent lines that fit the context window, including the stack trace, can be sent.

Key Characteristics of the Context Window

  • Hard limit: Requests that exceed the window are rejected by the API or silently truncated by the application.
  • Attention cost: Standard attention compares every token with every other token, so computation grows faster than the input length.
  • Position effects: Models often recall information at the start and end of a long input more reliably than information in the middle.
  • Latency and price: Longer inputs increase the time to the first generated token and the cost of each request.
  • Separate output cap: Many models also enforce a smaller, independent limit on the number of generated tokens.
  • Memory footprint: Serving long contexts requires a key-value cache whose size grows with every token held in the window.

How the Context Window Works

  1. Assembly: The application combines system instructions, conversation history, retrieved documents and the new user message into one prompt.
  2. Counting: A tokenizer calculates the token count of the assembled prompt before transmission.
  3. Reservation: A portion of the window is reserved for the expected length of the response.
  4. Fitting: If the prompt exceeds the remaining budget, the application drops, truncates or summarises the least important content, a practice known as context engineering.
  5. Processing: The model applies attention across every token inside the window, as described in transformer architecture.
  6. Generation: Generated tokens are appended to the window until the response completes or the token limit is reached.

Attention Cost and KV Cache Memory Formulas

Two costs grow with the number of tokens in the window: attention computation grows with the square of that number, and the memory of the key-value (KV) cache grows in direct proportion to it.

Attention Cost Grows with n Squared

  • : the number of tokens currently inside the context window.
  • : the query and key matrices, each containing one row per token.
  • : the dimension of each query and key vector.
  • : the score matrix, containing one row and one column for every token.

In words, every token is compared with every other token, so doubling the input length quadruples the number of attention scores. A causal mask removes approximately half of these comparisons, but the growth remains quadratic.

KV Cache Size Formula

During generation, the model stores the key and value vectors of every previous token, so it avoids recalculating them for each newly generated token.

  • : the factor for storing one key vector and one value vector.
  • : the number of transformer layers in the model.
  • : the number of heads that store keys and values, which equals every head in standard attention and fewer in multi-query or grouped-query attention.
  • : the dimension of each attention head.
  • : the number of tokens currently held in the cache during generation.
  • : the number of bytes for each stored value, which is 2 for 16-bit precision.

The cache therefore grows linearly, because every additional token contributes an identical, predetermined amount of accelerator memory to each individual request.

Worked Example with Round Model Sizes

The sizes below are round, illustrative numbers chosen for convenient arithmetic, rather than the published specification of any particular model.

  1. Per-token memory: Every cached token requires half a mebibyte of memory across all layers and attention heads of this illustrative configuration.
  2. Stack trace and source file: A prompt containing a long traceback and its related module needs almost four gibibytes of cache.
  3. Repository-scale prompt: Sixteen times as many tokens require sixteen times the memory, while the score matrix grows by a factor of 256.
  4. Fewer key-value heads: Sharing keys and values across heads reduces the cache in proportion to the number of heads removed.
Python
# Round, illustrative model sizes (not a specific real model)
LAYERS = 32
KV_HEADS = 32     # heads that store keys and values
HEAD_DIM = 128
BYTES = 2         # 16-bit numbers

def kv_cache_bytes(tokens, kv_heads=KV_HEADS):
    # 2 (keys and values) x layers x heads x head_dim x tokens x bytes
    return 2 * LAYERS * kv_heads * HEAD_DIM * tokens * BYTES

GIB = 1024 ** 3
print(f"Per token: {kv_cache_bytes(1) // 1024} KiB")
print("tokens    KV cache   scores per head and layer")
for n in (1_000, 8_000, 32_000, 128_000):
    print(f"{n:7,}  {kv_cache_bytes(n) / GIB:6.2f} GiB   {n * n:>17,}")

# Fewer key-value heads shrink the cache by the same factor
n = 32_000
for heads in (32, 8, 1):
    print(f"KV heads = {heads:2}, {n:,} tokens: {kv_cache_bytes(n, heads) / GIB:6.2f} GiB")
Output
Per token: 512 KiB
tokens    KV cache   scores per head and layer
  1,000    0.49 GiB           1,000,000
  8,000    3.91 GiB          64,000,000
 32,000   15.62 GiB       1,024,000,000
128,000   62.50 GiB      16,384,000,000
KV heads = 32, 32,000 tokens:  15.62 GiB
KV heads =  8, 32,000 tokens:   3.91 GiB
KV heads =  1, 32,000 tokens:   0.49 GiB
KV cache memory grows linearly with context length, while the number of attention scores grows with its squareTwo curves over context lengths from 0 to 128,000 tokens, each shown as a share of its own value at 128,000 tokens. KV cache memory is a straight line: 25 percent at 32,000 tokens (15.62 GiB) and 100 percent at 128,000 tokens (62.5 GiB) for an illustrative model with 32 layers, 32 KV heads, head dimension 128 and 16-bit values. Attention scores follow a parabola: only 6.25 percent at 32,000 tokens and 100 percent at 128,000 tokens. The model sizes are round illustrative numbers, not a real model.50%100%032k64k96k128ktokens in the context windowshare of value at 128kKV cache memory (linear)attention scores (quadratic)62.5 GiB of cache
KV cache memory grows linearly with context length, while the number of attention scores grows with its square
  • Memory is linear, computation is quadratic: Moving from 32,000 to 128,000 tokens multiplies the cache by four and the attention scores by sixteen.
  • Head sharing is the principal lever: Reducing the number of key-value heads shrinks the cache without shortening the context length.

These formulas explain why long prompts increase latency and price: computation grows with the square of the input, and cache memory limits how many long requests one accelerator can serve concurrently.

Example: Fitting a Log into the Context Window

The program below keeps the newest lines of a Python error log that fit a 60-token window after reserving 15 tokens for the answer.

Python
import re

def count_tokens(text):
    # Rough stand-in for a real tokenizer: words, numbers and symbols
    return len(re.findall(r"\w+|[^\w\s]", text))

WINDOW = 60           # total tokens the model accepts
RESERVED_OUTPUT = 15  # tokens set aside for the answer
instruction = "Explain the root cause of this Python traceback."

log = [f"INFO worker {n} processed batch {n}" for n in range(1, 9)] + [
    "Traceback (most recent call last):",
    'File "app.py", line 42, in load',
    "KeyError: 'user_id'",
]

def fit_to_window(instruction, lines):
    budget = WINDOW - RESERVED_OUTPUT - count_tokens(instruction)
    kept = []
    # Walk backwards so the newest lines, including the error, are kept
    for line in reversed(lines):
        cost = count_tokens(line)
        if cost > budget:
            break
        kept.insert(0, line)
        budget -= cost
    return kept, budget

total = count_tokens(instruction) + sum(count_tokens(l) for l in log)
kept, left = fit_to_window(instruction, log)
print(f"Full prompt: {total} tokens, window: {WINDOW}, reserved for output: {RESERVED_OUTPUT}")
print(f"Kept {len(kept)} of {len(log)} log lines, {left} tokens unused:")
for line in kept:
    print("  ", line)
Output
Full prompt: 82 tokens, window: 60, reserved for output: 15
Kept 4 of 11 log lines, 5 tokens unused:
   INFO worker 8 processed batch 8
   Traceback (most recent call last):
   File "app.py", line 42, in load
   KeyError: 'user_id'
  • Newest lines survive: Walking backwards preserves the traceback and the KeyError, which are the lines a model needs to diagnose the failure.
  • Output is budgeted first: Reserving 15 tokens prevents the prompt from consuming the space required for the explanation.
  • Production systems do more: Real assistants count with the model's own tokenizer and often summarise older lines instead of discarding them.

Applications of a Large Context Window

  • Whole-file code review: Reviewing an entire module or a complete pull request diff in a single request.
  • Log analysis: Examining lengthy stack traces and surrounding log entries during incident investigation.
  • Multi-turn assistants: Retaining earlier conversation turns so that follow-up questions are interpreted correctly.
  • Document question answering: Supplying API references directly, or selecting relevant passages through RAG.
  • Agent workflows: Accumulating tool results and intermediate reasoning steps for an AI agent.
  • Few-shot prompting: Including several worked examples that demonstrate the required output structure.

Advantages

  • Fewer retrieval steps: A large window can accommodate entire documents without a separate retrieval pipeline.
  • Better coherence: Longer history lets the model remain consistent across an extended conversation or a multi-file codebase.
  • Richer instructions: Detailed specifications, coding conventions and examples fit alongside the primary task.
  • Simpler engineering: Moderately sized inputs require less chunking, ranking and summarisation logic.

Limitations

  • Higher cost: Every token in the window is processed and billed on every request, including repeated conversation history.
  • Increased latency: Lengthy prompts increase the latency before the first generated token appears.
  • Degraded recall: Relevant details positioned in the middle of very long inputs are more likely to be overlooked.
  • Distracting content: Loosely related text can mislead the model and increase LLM hallucination.
  • No persistence: Content leaves the model's view once it is dropped from the window, so long-term memory requires external storage.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In what unit is a context window measured?

Frequently Asked Questions

What happens when a prompt exceeds the context window?

The API usually rejects the request with an error, or the application truncates the input before sending it. Any content that is removed is invisible to the model, so it cannot influence the response.

Does the context window include the model's output?

Yes. Input and generated output share the same token limit, so a very long prompt reduces the space available for the answer. Many applications reserve a fixed number of tokens for output.

Is a bigger context window always better?

Not always. Longer inputs cost more, respond more slowly and can bury important details, so a smaller set of relevant content often produces better answers.

How is the context window different from memory?

The context window holds only what is sent in the current request. Persistent memory requires the application to store information externally and add the relevant parts to later prompts.