Multi-Agent Systems
Multi-agent systems are AI architectures in which several specialised agents, each with its own instructions, tools and context, cooperate to complete one goal. A coordination mechanism, such as an orchestrator, handoffs or shared state, decides which agent acts next and how their results are combined.
- Specialisation: Each AI agent has a narrow role, such as log analysis, and only the tools that role needs.
- Orchestrator-worker: In the orchestrator-worker pattern, a central agent decomposes the goal, dispatches sub-tasks and merges the results.
- Supervisor: A managing agent reviews worker output and decides whether to accept it, retry or escalate.
- Handoffs: Agent handoffs transfer control and context from one agent to another when the task leaves its role.
- Shared state: A common store, often called a blackboard, holds findings that every agent can read and write.
For example, an orchestrator investigating a 5xx spike on the checkout service dispatches log, metrics and deploy agents in parallel and merges their findings into one timeline.
Key Characteristics of Multi-Agent Systems
- Decomposition: A goal is divided into sub-tasks that correspond to agent roles.
- Isolated context: Each agent receives only the data for its sub-task, which keeps prompts short and focused.
- Least privilege: Tool permissions are scoped per agent, so the log agent cannot trigger a rollback.
- Parallelism: Independent sub-tasks run concurrently, which reduces end-to-end latency.
- Communication protocol: Messages, handoffs or shared state define how agents exchange results, and A2A standardises this between separate systems.
- Aggregation: A merge step reconciles findings, resolves conflicts and produces a single decision.
Coordination Patterns
| Pattern | Who controls the flow | Incident example |
|---|---|---|
| Orchestrator-worker | A central agent plans and dispatches | Log, metrics and deploy sub-tasks sent in parallel |
| Supervisor | A manager reviews and may reject output | A summary without evidence is returned for rework |
| Handoffs | Control passes from agent to agent | Triage hands a database alert to a database agent |
| Shared state | Agents read and write a common store | Every finding lands on one incident timeline |
These four patterns are the main forms of multi-agent orchestration, and they extend the multi-agent collaboration pattern described in agentic design patterns.
How Multi-Agent Systems Work
- Goal intake: An alert or request reaches the orchestrator, together with policies such as approval rules.
- Decomposition: The orchestrator splits the goal into sub-tasks and assigns an agent to each.
- Dispatch: Sub-tasks go to worker agents, in parallel wherever they are independent.
- Execution: Each worker runs its own loop through tool calling and returns a structured finding.
- State update: Findings are written to shared state with timestamps and sources.
- Merge and decision: The orchestrator correlates findings, resolves conflicts and selects the next action.
- Handoff or approval: Work passes to another agent or to a person, as in human-in-the-loop approval of a production rollback.
Example: An Orchestrator with Three Sub-Agents
The program below dispatches log, metrics and deploy sub-agents in parallel, merges their findings into a timeline and applies the rollback approval rule.
# Orchestrator-worker: one orchestrator, three specialist sub-agents.
# Standard library only; fixed data and rules stand in for tools and LLMs.
from concurrent.futures import ThreadPoolExecutor
LOGS = ["14:05 502 payment client timeout", "14:06 502 payment client timeout"]
METRICS = {"5xx_rate": 12.0, "baseline": 0.2, "spike_start": "14:05"}
DEPLOYS = [{"service": "checkout", "version": "v2.4.1", "time": "14:03"}]
def log_agent(task):
errors = [line for line in LOGS if "502" in line]
return {"agent": "logs", "finding": f"{len(errors)} x 502 payment client timeout",
"since": errors[0].split()[0]}
def metrics_agent(task):
ratio = METRICS["5xx_rate"] / METRICS["baseline"]
return {"agent": "metrics", "finding": f"5xx rate {METRICS['5xx_rate']}% ({ratio:.0f}x baseline)",
"since": METRICS["spike_start"]}
def deploy_agent(task):
d = DEPLOYS[-1]
return {"agent": "deploys", "finding": f"{d['service']} {d['version']} deployed",
"since": d["time"], "version": d["version"]}
WORKERS = {"logs": log_agent, "metrics": metrics_agent, "deploys": deploy_agent}
def orchestrate(goal):
shared_state = {"goal": goal, "findings": []} # blackboard all agents write to
with ThreadPoolExecutor() as pool: # dispatch the sub-agents in parallel
results = pool.map(lambda name: WORKERS[name](goal), WORKERS)
shared_state["findings"].extend(results)
for f in sorted(shared_state["findings"], key=lambda f: f["since"]): # merged timeline
print(f" {f['since']} [{f['agent']}] {f['finding']}")
deploy = next(f for f in shared_state["findings"] if f["agent"] == "deploys")
spike = min(f["since"] for f in shared_state["findings"] if f["agent"] != "deploys")
if deploy["since"] < spike: # deploy came first: likely cause
return f"rollback of {deploy['version']} proposed, held for human approval"
return "no deploy before the spike, handing off to the on-call engineer"
print("Goal: investigate the 5xx spike on checkout")
print("Decision: " + orchestrate("investigate the 5xx spike on checkout"))Output:
Goal: investigate the 5xx spike on checkout
14:03 [deploys] checkout v2.4.1 deployed
14:05 [logs] 2 x 502 payment client timeout
14:05 [metrics] 5xx rate 12.0% (60x baseline)
Decision: rollback of v2.4.1 proposed, held for human approval- Parallel dispatch: The thread pool runs the three sub-agents concurrently, and pool.map returns results in dispatch order, so the output is deterministic.
- Structured findings: Every agent returns the same fields, which lets the orchestrator sort them into one timeline.
- Correlation and gate: The deploy at 14:03 precedes the spike at 14:05, so the orchestrator proposes a rollback and holds it for approval.
Applications of Multi-Agent Systems
- Incident response: Specialist agents for logs, metrics, traces and deployments.
- Software development: Planner, coder and reviewer agents collaborating on one pull request.
- Research and analysis: Parallel search agents feeding a synthesiser that writes the report.
- Security operations: Enrichment, correlation and response agents for alert triage.
- Data engineering: Agents that validate, transform and document individual pipeline stages.
Advantages
- Focused prompts: Narrow roles reduce conflicting instructions and context length per agent.
- Lower latency: Parallel workers complete independent checks sooner than one sequential agent.
- Safer permissions: Scoped tool access limits the damage a single faulty agent can cause.
- Modularity: Agents can be tested, replaced or upgraded individually.
Limitations
- Coordination overhead: Additional model calls and messages raise cost and add failure points.
- Error propagation: One incorrect finding can mislead the merge step and every later decision.
- Conflicting results: Agents can disagree, which requires explicit conflict resolution rules.
- Harder debugging: Failures can span several agents, so AI agent evaluation must trace messages across them.
- Security surface: Every message between agents is a possible channel for prompt injection.
Frameworks such as LangGraph and CrewAI implement these patterns, as compared in LangGraph vs CrewAI vs AutoGen.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the orchestrator-worker pattern, which component merges the sub-agent results?
Frequently Asked Questions
What is a multi-agent system in AI?
It is an architecture in which several specialised AI agents, each with its own instructions and tools, cooperate on one goal. A coordination mechanism such as an orchestrator, handoffs or shared state decides who acts and how results are combined.
When should a multi-agent system be used instead of a single agent?
Use one when a single agent's context or tool set grows too large, when sub-tasks can run in parallel, or when permissions must be separated by role. For small, bounded tasks a single agent is simpler and cheaper.
What is the difference between an orchestrator and a supervisor agent?
An orchestrator plans the work, dispatches sub-tasks and merges results. A supervisor focuses on oversight, reviewing each worker's output and deciding whether to accept it, request a retry or escalate to a person.
What is a handoff in a multi-agent system?
A handoff transfers control of a task, together with its context, from one agent to another. For example, a triage agent hands a database alert to a database specialist agent instead of handling it itself.
Related Articles
- Agentic Design PatternsAgentic design patterns explained: reflection, tool use, planning, multi-agent collaboration, routing and evaluator-optimizer, with Python reflection code.
- Agentic AI vs AI AgentsAgentic AI vs AI agents explained: a system-level approach versus the individual worker, with a comparison table, autonomy levels and an incident example.
- What is an AI AgentAn AI agent explained: its definition, key characteristics, how the perceive, decide and act loop works, a Python CI fixing agent, uses and limitations.
- Human-in-the-Loop in Agentic AIHuman-in-the-loop in agentic AI: risk-based approval gates, escalation and audit logs, with a Python approval queue for an incident rollback and limits.
- A2A Protocol vs MCPA2A protocol vs MCP: agent-to-agent delegation versus agent-to-tool calls, a comparison table, when to use each, and an incident example that uses both.
- LangGraph vs CrewAI vs AutoGenLangGraph vs CrewAI vs AutoGen compared: control model, state, human-in-the-loop, multi-agent style and learning curve, with one incident agent in each.