Agentic Design Patterns
Agentic design patterns are reusable structures for arranging language model calls, tools and control logic so that an AI system completes multi-step tasks reliably. Each pattern solves one recurring problem in agentic AI, and production systems usually combine several of them.
- Reflection: The model critiques its own output against explicit criteria and revises it before returning it.
- Tool use: The model calls functions or APIs to read live data and act on external systems.
- Planning: The model decomposes a goal into ordered sub-tasks before executing them.
- Multi-agent collaboration: Specialist agents with separate prompts and tools divide the work.
- Routing: A classification step sends each input to the most suitable handler.
- Evaluator-optimizer: A separate evaluator scores each output, and a generator revises it until the score passes.
For example, an incident response system routes a 5xx alert to the checkout triage agent, which plans its checks, calls tools and reflects on its summary before posting it.
Key Characteristics of Agentic Design Patterns
- Composability: Patterns nest inside each other, such as reflection inside one worker of a multi-agent system.
- Model independence: Each pattern describes control flow, so it applies to any capable LLM or framework.
- Explicit termination: Every loop needs a stopping condition, such as a round limit or a passing score.
- Cost trade-off: Most patterns add inference calls, exchanging latency and cost for accuracy or coverage.
- Observability: Each step yields an inspectable intermediate result, which simplifies tracing and debugging.
Agentic Design Patterns Compared
| Pattern | Problem it solves | Incident example |
|---|---|---|
| Reflection | Omissions and errors in a single draft | Checking an incident summary against a checklist |
| Tool use | No access to live data | Querying logs, metrics and deploy history |
| Planning | Goals too broad for one step | Ordering the checks before running them |
| Multi-agent | One context or tool set becomes too large | Separate log, metrics and deploy agents |
| Routing | Mixed inputs need different handling | Sending checkout alerts to the payments runbook |
| Evaluator-optimizer | Quality must meet a measurable bar | Revising a rollback plan until a policy check passes |
How Agentic Design Patterns Work
AI agent design patterns are usually added incrementally, starting from the simplest design that meets the requirement.
- Single call: Begin with one prompt, for example to summarise an error log, and record where it fails.
- Tool use: Expose read-only tools such as a log search through tool calling when the model needs live data.
- Routing: Classify incoming alerts by service or severity, and send each class to a specialised prompt or agent.
- Planning: For open-ended goals, have the model produce an ordered plan first, as described in planning in AI agents.
- Reflection or evaluation: When quality can be checked against criteria, insert a critique and revision loop.
- Multiple agents: Divide the work only when one context grows too large, as covered in multi-agent systems.
- Bounds: Add round limits, timeouts and human-in-the-loop gates for irreversible actions.
The ReAct loop combines tool use with step-by-step reasoning. Published agent-building guidance from model providers also names routing, orchestrator-workers and evaluator-optimizer as core agentic workflow patterns.
Example: The Reflection Pattern for an Incident Summary
The program below drafts a checkout incident summary, critiques it against a checklist and revises it until every requirement passes. Fixed rules stand in for the LLM calls.
# Reflection: draft, critique against a checklist, revise, repeat.
# Standard library only; fixed rules stand in for the LLM calls.
EVIDENCE = {
"impact": "Impact: 5xx rate 12% (normal 0.2%) since 14:05.",
"cause": "Suspected cause: deploy v2.4.1 at 14:03.",
"action": "Action: rollback of v2.4.1 awaiting human approval.",
"owner": "Owner: on-call engineer for checkout.",
}
CHECKLIST = [ # (requirement, test, evidence key used to repair it)
("names the affected service", lambda t: "checkout" in t.lower(), None),
("quantifies impact with a metric", lambda t: "%" in t, "impact"),
("states the suspected cause", lambda t: "v2.4.1" in t, "cause"),
("states the action and approval status", lambda t: "approval" in t, "action"),
("names an owner", lambda t: "owner" in t.lower(), "owner"),
]
def critique(summary):
# Critic step: return every checklist item the summary fails.
return [(req, key) for req, test, key in CHECKLIST if not test(summary)]
def revise(summary, problems, per_round=2):
# Reviser step: fix at most two problems per round, as one LLM rewrite might.
for _, key in problems[:per_round]:
summary += " " + EVIDENCE[key]
return summary
summary = "Checkout is failing for some users." # first draft
for round_no in range(1, 5): # hard cap on reflection rounds
problems = critique(summary)
print(f"Round {round_no}: {len(problems)} problem(s)")
for req, _ in problems:
print(f" missing: {req}")
if not problems:
break
summary = revise(summary, problems)
print("Final summary:")
for sentence in summary.split(". "):
print(" " + sentence.rstrip(".") + ".")Output:
Round 1: 4 problem(s)
missing: quantifies impact with a metric
missing: states the suspected cause
missing: states the action and approval status
missing: names an owner
Round 2: 2 problem(s)
missing: states the action and approval status
missing: names an owner
Round 3: 0 problem(s)
Final summary:
Checkout is failing for some users.
Impact: 5xx rate 12% (normal 0.2%) since 14:05.
Suspected cause: deploy v2.4.1 at 14:03.
Action: rollback of v2.4.1 awaiting human approval.
Owner: on-call engineer for checkout.- Explicit criteria: The checklist makes the critique specific, so each round reports exactly which requirements are missing.
- Convergence: The reviser fixes two problems per round, so the critique passes at round 3, well inside the cap of 4.
- Evaluator-optimizer variant: Moving the checklist into a separate evaluator model or automated test turns this loop into the evaluator-optimizer pattern.
Applications of Agentic Design Patterns
- Incident reporting: Reflection checks postmortem drafts against a required template.
- Code generation: Evaluator-optimizer loops run unit tests and return failures to the model.
- Alert triage: The routing pattern sends alerts to database, network or application specialists.
- Research tasks: Planning splits a question into several searches, as in agentic RAG.
- Data extraction: Reflection validates extracted JSON against a schema before storage.
Advantages
- Reliability: Critique and evaluation steps catch omissions before output reaches a person.
- Shared vocabulary: Named patterns make agent designs easier to discuss and review.
- Incremental adoption: Patterns are added one at a time, when measured failures justify them.
- Portability: The same pattern maps onto LangGraph, CrewAI or plain Python.
Limitations
- Latency: Reflection and evaluation multiply the number of model calls per task.
- Self-assessment bias: A model critiquing its own output can overlook errors it introduced in the draft.
- Over-engineering: Multi-agent designs add coordination overhead that simple tasks do not need.
- Runaway loops: Without a round cap, a critic that never passes causes unbounded cost.
- Testing effort: Combined patterns create many execution paths, which require AI agent evaluation.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. Which pattern sends each incoming alert to the most suitable handler?
Frequently Asked Questions
What are the main agentic design patterns?
The most widely cited are reflection, tool use, planning and multi-agent collaboration. Routing and evaluator-optimizer are also common, and most production systems combine several of them.
What is the difference between reflection and evaluator-optimizer?
In reflection, the same model critiques and revises its own output. In evaluator-optimizer, a separate evaluator, which can be another model or an automated test, scores the output and the generator revises it until the score passes.
Which agentic design pattern should be used first?
Start with a single model call plus tool use, then measure where it fails. Add routing, planning or reflection only when a specific failure justifies the extra latency and cost.
Do agentic design patterns need a framework?
No. Each pattern is a control-flow structure that can be written in plain Python. Frameworks such as LangGraph or CrewAI provide ready components for loops, state and multiple agents, which saves work in larger systems.
Related Articles
- What is Agentic AIAgentic AI explained: what it is, its key characteristics, how the agent loop plans and uses tools, a Python incident agent example, uses and limitations.
- Multi-Agent SystemsMulti-agent systems explained: orchestrator-worker, supervisor, handoffs and shared state, with a Python orchestrator that merges three sub-agent findings.
- Planning and Reasoning in AI AgentsLearn planning and reasoning in AI agents: task decomposition, dependencies, plan-and-execute and replanning, with a Python flaky test fix example.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- ReAct Agent ExplainedLearn how a ReAct agent alternates thought, action and observation, with a Python example that traces a failing unit test to a commit and opens a ticket.
- AI Agent EvaluationAI agent evaluation explained: task success, tool-call accuracy, trajectory checks, cost, latency and regression suites, with a Python trajectory scorer.