What is Context Engineering
Context engineering is the practice of deciding which information enters a language model's context window for each request, including instructions, retrieved documents, conversation history, tool results and examples. It extends prompt engineering from wording a single prompt to selecting, ranking, compressing and ordering everything the model reads within a fixed token budget.
- Context: Everything the model receives in one request, which is the only information it can use beyond its training data.
- Token budget: The share of the context window available for inputs after space is reserved for the response.
- Selection: Retrieval components choose candidate snippets, such as log entries, runbook sections or code files, for the current task.
- Compression: Long material is summarised or truncated so that more relevant information fits, a core part of context window optimisation.
- Assembly: The final context is ordered and labelled, so the model can distinguish instructions from data and identify each source.
For example, an error log summariser includes the failing log lines, the relevant runbook entry and the latest deploy note, and excludes older wiki pages that would consume the budget.
Key Characteristics of Context Engineering
- Dynamic: The context is rebuilt for every request, so LLM context management is a per-request task rather than a one-time setup.
- Budget-constrained: Every snippet competes for limited tokens, so including one item can mean excluding another.
- Relevance-ranked: Candidates are ordered by a relevance score, usually produced by a retriever and sometimes refined by reranking.
- Source-aware: Each snippet carries a label, such as log, runbook or deploy, so the model can cite and weigh its origin.
- Layered: A stable system prompt sits alongside variable retrieved data, recent history and the user's question.
- Measurable: Answer quality, latency and cost can be compared across different context strategies on the same test questions.
How Context Engineering Works
- Reserve space: Subtract the system prompt, the question and the maximum response length from the context window to obtain the input budget.
- Gather candidates: Collect possible context from retrieval, conversation memory, tool outputs and configuration data.
- Score and rank: Order the candidates by relevance to the current question, using embeddings, keywords or recency.
- Fit the budget: Add candidates in rank order while the estimated token count stays within the budget, and record what was excluded.
- Compress where necessary: Summarise long histories or documents that are important but too large to include in full.
- Assemble and label: Place the selected items in a consistent order, with clear delimiters between instructions and untrusted data.
Example: Assembling a Context Window in Python
The program below selects ranked snippets for an incident question until an illustrative token budget is exhausted.
# Assemble a context window from ranked snippets under a token budget
BUDGET = 60 # tokens reserved for retrieved context (illustrative)
def estimate_tokens(text):
return len(text) // 4 + 1 # rough rule: about 4 characters per token
# (relevance score, source, text), already ranked by a retriever
snippets = [
(0.92, "log", "10:02:11 ERROR payments-db connection refused on port 5432"),
(0.88, "runbook", "If payments-db refuses connections, check the pool limit and restart the pgbouncer pod."),
(0.71, "deploy", "Deploy 4812 at 09:58 changed DB_POOL_SIZE from 50 to 5 in checkout-api."),
(0.40, "wiki", "The payments team was formed in 2021 and owns billing, refunds and invoices."),
(0.35, "log", "10:01:00 INFO web-frontend health check ok"),
]
context, used, dropped = [], 0, []
for score, source, text in sorted(snippets, reverse=True):
cost = estimate_tokens(text)
if used + cost <= BUDGET:
context.append(f"[{source}] {text}")
used += cost
else:
dropped.append(f"{source} ({cost} tokens, score {score})")
print(f"Used {used} of {BUDGET} tokens")
print("Dropped:", ", ".join(dropped))
print("--- context block ---")
print("\n".join(context))Used 55 of 60 tokens
Dropped: wiki (20 tokens, score 0.4), log (11 tokens, score 0.35)
--- context block ---
[log] 10:02:11 ERROR payments-db connection refused on port 5432
[runbook] If payments-db refuses connections, check the pool limit and restart the pgbouncer pod.
[deploy] Deploy 4812 at 09:58 changed DB_POOL_SIZE from 50 to 5 in checkout-api.- Greedy selection: The three highest-ranked snippets fit within 55 tokens, and the two lower-ranked snippets are excluded because they would exceed the budget.
- Approximate counting: The four-characters-per-token rule is only an estimate, and production systems count tokens with the model's own tokenizer.
- Traceable decisions: Recording the dropped snippets explains why the model did not mention them, which simplifies debugging when an answer is incomplete.
Applications of Context Engineering
- Retrieval systems: Choosing which passages RAG pipelines place in front of the model.
- Coding assistants: Selecting the relevant files, function signatures and test failures from a large repository.
- Agent memory: Deciding which past steps and observations remain visible, as discussed in memory in AI agents.
- Incident response: Combining recent logs, metrics, deploy history and runbooks for an on-call summary.
- Tool integration: Formatting results from external systems, including those connected through MCP.
Advantages
- Higher accuracy: Relevant, well-labelled context reduces guessing and unsupported answers.
- Lower cost: Excluding irrelevant material reduces input tokens and response latency.
- Scalability: Knowledge bases far larger than the context window become usable through selection.
- Debuggability: Logged context decisions reveal whether a failure came from retrieval or from generation.
Limitations
- Retrieval errors: An incorrectly ranked or missing snippet cannot be recovered by the model.
- Compression loss: Summaries can omit the one detail that the question depends on.
- Positional sensitivity: Models can pay less attention to information placed in the middle of a long context.
- Security exposure: Retrieved documents can contain hidden instructions, a risk described in prompt injection.
- Engineering overhead: Ranking, token counting and evaluation add components that require maintenance and monitoring.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does context engineering decide?
Frequently Asked Questions
Context engineering vs prompt engineering: what is the difference?
Prompt engineering focuses on how the instructions are worded. Context engineering covers everything the model receives, including retrieved documents, history and tool results, and decides what to include within the token budget.
Why not put all available information into a large context window?
More input raises cost and latency, and irrelevant material can distract the model from the facts that matter. Selecting a smaller, relevant context usually produces better answers.
Is context engineering the same as RAG?
No. RAG is one source of context, the retrieval of documents for a question. Context engineering also manages instructions, conversation history, tool results and the order in which everything appears.
How are tokens counted when building a context?
Production systems use the tokenizer that belongs to the target model, because each model splits text differently. Character-based rules give only a rough estimate for planning.
Related Articles
- What is Prompt EngineeringLearn what prompt engineering is, how to design a prompt step by step, its key characteristics, uses and limits, with a Python code review prompt example.
- Context Window in LLMLearn what the context window of an LLM is, how tokens fill it, what happens past the limit and how to fit long logs, with a Python truncation example.
- What is RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) explained: how it retrieves document passages to ground LLM answers, its key steps, a Python example and its limits.
- Memory in AI AgentsLearn how memory in AI agents works: context window buffers, summarisation and long-term stores, with a Python example that recalls past DevOps incidents.
- System Prompt vs User PromptSystem prompt vs user prompt: who writes each, what belongs in each, how models prioritise them, with a comparison table and a Python code review example.
- Reranking in RAGReranking in RAG explained: how a second, stricter scorer reorders retrieved passages, cross-encoders vs bi-encoders, a Python example and the trade-offs.