…
Skip to content
Topics
On this page

CrewAI Tutorial

CrewAI is an open-source Python framework for building a team of language model agents, where each agent has a role, a goal and a set of tools. A crew runs a list of tasks in a defined process and passes each result forward, so specialised agents can complete a multi-step job together.

  • Agent: A model with a role, a goal, a short backstory and the tools it may call.
  • Task: One unit of work with a description, an expected output and the agent that owns it.
  • Crew: The object that holds the agents and tasks and starts the run with kickoff().
  • Process: The order of execution; sequential runs tasks in list order, hierarchical adds a manager agent.
  • Tool: A Python function that an agent can call, such as reading logs or listing deploys.
  • Context: A link between tasks, so a later task receives the outputs of earlier tasks.
A CrewAI triage crew with a sequential processThree CrewAI agents run in a sequential process. The log analyst calls read_logs and get_metrics, the deploy analyst calls list_recent_deploys, and the incident lead writes the summary without tools. The recommendation then goes to an approval gate where the on-call engineer decides, and only an approved rollback reaches the deploy system. The flow is illustrative.Process.sequentialLog analystread_logs, get_metricsDeploy analystlist_recent_deploysIncident leadsummary, no toolsApproval gateon-call engineerRollbackdeploy system
A CrewAI triage crew with a sequential process

For example, a crew of three agents triages a 5xx spike on the checkout service: a log analyst reads logs and metrics, a deploy analyst checks recent deploys, and an incident lead writes the summary.

The CrewAI Python example below uses three read-only tools, read_logs, get_metrics and list_recent_deploys, and the same agent is also built with LangGraph and n8n so the frameworks can be compared directly. In a CrewAI multi-agent design, the work is divided by role, and a production rollback stays outside the crew behind a human-in-the-loop approval. This division makes each agent's responsibility explicit, although it also adds model calls compared with a single AI agent.

Prerequisites

  • Runtime: CrewAI supports recent Python 3 releases, and the examples were tested with Python 3.11.
  • Background: Familiarity with tool calling in LLMs and the idea of multi-agent systems.
  • Credentials: An Anthropic API key (or another provider key), stored in an environment variable.
  • Comparison: The LangGraph tutorial builds the same agent as a graph, which is useful for comparison.

Setup

Create a project folder and a virtual environment, then install the pinned CrewAI release.

Bash
mkdir triage-crew && cd triage-crew
python3.11 -m venv .venv
source .venv/bin/activate
pip install crewai==1.0.0
export ANTHROPIC_API_KEY="paste-your-key-here"
export CREW_MODEL="anthropic/claude-sonnet-4-5"

Step 1: Write the Incident Tools as Plain Functions

The three tools start as ordinary functions over canned data, so they can be tested without a model or a network. In production, each function would call the logging service, the metrics API and the deploy system.

Python
# incident_tools.py: canned data for the three triage tools.
LOGS = {
    "checkout": [
        "14:02:11 ERROR POST /checkout 502 upstream payments-api timeout after 3000 ms",
        "14:02:14 ERROR POST /checkout 502 upstream payments-api timeout after 3000 ms",
        "14:02:19 WARN  retry budget exhausted for payments-api",
    ],
}
METRICS = {
    "checkout": {"5xx_rate_pct": {"13:55": 0.2, "14:00": 0.3, "14:05": 7.9}, "p95_latency_ms": {"13:55": 310, "14:05": 3020}},
}
DEPLOYS = [
    {"service": "checkout", "version": "v2.41.0", "time": "14:00", "change": "payments client timeout 10 s -> 3 s"},
    {"service": "search", "version": "v1.9.2", "time": "12:30", "change": "new ranking weights"},
]

def read_logs(service: str, limit: int = 3) -> str:
    return "\n".join(LOGS.get(service, [])[:limit]) or "no log lines"

def get_metrics(service: str) -> str:
    m = METRICS.get(service, {})
    return "; ".join(f"{name}: {series}" for name, series in m.items()) or "no metrics"

def list_recent_deploys(service: str) -> str:
    rows = [d for d in DEPLOYS if d["service"] == service]
    return "\n".join(f"{d['time']} {d['version']}: {d['change']}" for d in rows) or "no deploys"

if __name__ == "__main__":
    print(read_logs("checkout"))
    print(get_metrics("checkout"))
    print(list_recent_deploys("checkout"))

Output:

Example
14:02:11 ERROR POST /checkout 502 upstream payments-api timeout after 3000 ms
14:02:14 ERROR POST /checkout 502 upstream payments-api timeout after 3000 ms
14:02:19 WARN  retry budget exhausted for payments-api
5xx_rate_pct: {'13:55': 0.2, '14:00': 0.3, '14:05': 7.9}; p95_latency_ms: {'13:55': 310, '14:05': 3020}
14:00 v2.41.0: payments client timeout 10 s -> 3 s
  • Safety: All three functions only read data, so an agent that calls them cannot change production.
  • Formatting: Each tool returns short text, which keeps the model's context small and easy to inspect.

Step 2: Define the CrewAI Agents and Tasks

Save this as triage_crew.py next to incident_tools.py. The @tool decorator turns each function into a CrewAI tool, and the docstring becomes the description the model reads.

Python
# triage_crew.py: three agents, three tasks, one sequential crew.
import os

from crewai import LLM, Agent, Crew, Process, Task
from crewai.tools import tool

import incident_tools as data

llm = LLM(model=os.environ["CREW_MODEL"], temperature=0)  # key read from ANTHROPIC_API_KEY

@tool("read_logs")
def read_logs(service: str) -> str:
    """Return the latest error log lines for a service."""
    return data.read_logs(service)

@tool("get_metrics")
def get_metrics(service: str) -> str:
    """Return the 5xx rate and p95 latency series for a service."""
    return data.get_metrics(service)

@tool("list_recent_deploys")
def list_recent_deploys(service: str) -> str:
    """Return recent deploys of a service with time, version and change."""
    return data.list_recent_deploys(service)

log_analyst = Agent(
    role="Log analyst",
    goal="Find the error pattern behind the 5xx spike on {service}",
    backstory="An SRE who reads logs and metrics and names the failing dependency.",
    tools=[read_logs, get_metrics], llm=llm, allow_delegation=False,
)
deploy_analyst = Agent(
    role="Deploy analyst",
    goal="Decide whether a recent deploy of {service} explains the spike",
    backstory="A release engineer who matches deploy times against incident timelines.",
    tools=[list_recent_deploys], llm=llm, allow_delegation=False,
)
incident_lead = Agent(
    role="Incident lead",
    goal="Write a short triage summary with one recommended action",
    backstory="An on-call lead who recommends actions but never performs them.",
    llm=llm, allow_delegation=False,
)
  • Templating: The {service} placeholder is filled from the inputs at kickoff, so one crew can triage any service.
  • Permissions: Each agent receives only the tools its role needs, and the incident lead receives none.
  • Descriptions: The model selects tools from their names and docstrings, so precise descriptions reduce incorrect or repeated tool calls.
  • Determinism: A temperature of 0 reduces variation between runs, which makes triage summaries easier to compare during evaluation.
  • Delegation: Setting allow_delegation=False prevents an agent from passing its work to another agent, which keeps the execution order predictable.

Step 3: Assemble the Crew with the CrewAI Sequential Process

Append the tasks and the crew to triage_crew.py, where the context list passes earlier task outputs forward to later tasks.

Python
logs_task = Task(
    description="Read logs and metrics for {service}. Report the main error, when it started and the 5xx rate.",
    expected_output="Three bullets: error pattern, start time, 5xx rate.",
    agent=log_analyst,
)
deploy_task = Task(
    description="List recent deploys of {service} and compare their times with the spike start.",
    expected_output="The deploy most likely to explain the spike, with one reason.",
    agent=deploy_analyst, context=[logs_task],
)
summary_task = Task(
    description=("Write the triage summary. End with one line: 'Recommended action: ROLLBACK <version>' "
                 "if a deploy explains the spike, otherwise 'Recommended action: INVESTIGATE'."),
    expected_output="Likely cause, evidence and the recommended action line.",
    agent=incident_lead, context=[logs_task, deploy_task],
)

crew = Crew(
    agents=[log_analyst, deploy_analyst, incident_lead],
    tasks=[logs_task, deploy_task, summary_task],
    process=Process.sequential, verbose=True,
)

if __name__ == "__main__":
    result = crew.kickoff(inputs={"service": "checkout"})
    print(result.raw)
  • Sequencing: CrewAI executes the three tasks in list order, and each task receives the outputs named in its context list as additional prompt content.
  • Specification: The expected_output text becomes part of the agent's instructions, so it acts as a lightweight contract for every task.
  • Hierarchy: Process.hierarchical adds a manager model that assigns tasks and validates results, which increases flexibility but also latency and token usage.
  • Logging: verbose=True prints every agent's reasoning and tool call, which is useful for debugging but should be redirected to structured logs in production.

Step 4: Gate the Rollback Behind Human Approval

The crew only recommends an action, and a separate function reads the last line of the summary and allows the rollback only after a named person approves it; in production, the approval comes from a chat button or input().

Python
# approval_gate.py: the rollback runs only after a named person approves it.
import re

SUMMARY = """Likely cause: deploy v2.41.0 of checkout at 14:00 cut the payments client timeout from 10 s to 3 s.
Evidence: 502 errors on POST /checkout start at 14:02; 5xx rate rose from 0.3% to 7.9%.
Recommended action: ROLLBACK v2.41.0"""

def proposed_rollback(summary):
    match = re.search(r"Recommended action: ROLLBACK (v[\d.]+)", summary)
    return match.group(1) if match else None

def approval_gate(summary, approver, approved):
    version = proposed_rollback(summary)
    if version is None:
        return "No rollback proposed. Summary posted to the incident channel."
    if not approved:
        return f"Rollback of {version} rejected by {approver}. No production change made."
    return f"Rollback of checkout {version} approved by {approver}. Calling the deploy system."

print(approval_gate(SUMMARY, "priya.oncall", approved=False))
print(approval_gate(SUMMARY, "priya.oncall", approved=True))
print(approval_gate("Recommended action: INVESTIGATE", "priya.oncall", approved=True))

Output:

Example
Rollback of v2.41.0 rejected by priya.oncall. No production change made.
Rollback of checkout v2.41.0 approved by priya.oncall. Calling the deploy system.
No rollback proposed. Summary posted to the incident channel.
  • Parsing: The summary task asks for one exact action line, so plain code can parse the recommendation reliably.
  • Alternatives: Setting human_input=True on a task asks for terminal feedback before the task finishes, which suits review but not a production gate.

Output

Run python triage_crew.py with the key exported in the same terminal session. Illustrative final output, where the model's wording varies between runs:

Example
Likely cause: deploy v2.41.0 of checkout at 14:00 reduced the payments client timeout from 10 s to 3 s.
Evidence: payments-api timeouts (502) begin at 14:02; 5xx rate rose from 0.3% to 7.9%; p95 latency reached 3020 ms.
Recommended action: ROLLBACK v2.41.0
  • Correlation: The deploy analyst connects the 14:00 deploy with the first timeout at 14:02, which is the central evidence for the recommendation.
  • Traceability: Every claim in the summary refers to a tool result, so an engineer can verify each statement against the logs, metrics and deploy history.
  • Approval: The summary is passed to approval_gate() from Step 4, and nothing changes in production until the on-call engineer approves.

Common Errors

  • ModuleNotFoundError: No module named 'crewai': The virtual environment is inactive, so activate it and run pip install crewai==1.0.0 again.
  • KeyError: 'CREW_MODEL': The model variable is not exported in the current shell, so export it before the run.
  • AuthenticationError from the model provider: The API key is missing or invalid, so export a valid ANTHROPIC_API_KEY in the same terminal.
  • Missing template variable error at kickoff: A {service} placeholder has no value, so pass every placeholder in kickoff(inputs=...).
  • Agent stopped due to iteration limit or time limit: An agent kept calling tools without finishing, so tighten the task description or raise max_iter on the agent.

Next Steps

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the CrewAI triage crew, which setting makes the tasks run one after another?

Frequently Asked Questions

What is CrewAI used for?

CrewAI is used to build teams of language model agents that each hold one role and pass results to each other. Typical uses are research pipelines, report writing, triage and review workflows where the work splits into clear steps.

Is CrewAI built on LangChain?

CrewAI is a separate Python framework with its own agent, task and crew classes. It can call many model providers through one model interface, and a crew can use tools written as plain Python functions.

What is the difference between a sequential and a hierarchical process in CrewAI?

A sequential process runs the tasks in the order they are listed, and each task can read the output of earlier tasks. A hierarchical process adds a manager that assigns tasks to agents and checks their results.

Can a CrewAI agent perform a production rollback on its own?

It can if it is given a rollback tool, which is why the safer design gives no agent that tool. The crew only recommends the action, and ordinary code performs it after a person approves.

Which models can CrewAI use?

CrewAI can call hosted models from several providers and local models through one model setting. The provider key is read from an environment variable, so the same crew runs with different models by changing configuration only.