…
Skip to content
Topics
On this page

LangGraph Tutorial

LangGraph is an open-source Python and JavaScript library for building stateful agents and workflows as graphs of nodes connected by edges. It stores a shared state, routes between steps with conditional logic and saves checkpoints, so a run can loop, branch, pause for human input and resume later.

  • State: A typed dictionary that every node reads and updates, which functions as the working memory of the execution.
  • Nodes: Ordinary Python functions that receive the current state and return only the keys they modify.
  • Edges: Unconditional or conditional transitions that determine which node executes next.
  • Checkpointer: A persistence component that records the state after every step under a thread identifier.
  • Interrupts: The LangGraph interrupt() function suspends execution inside a node until an external caller supplies a value.
The LangGraph incident triage graphThe graph runs from START to gather_evidence, then to decide. A conditional edge sends a rollback decision to the approve node, which calls interrupt() and waits for the on-call engineer, and sends a monitor decision straight to report. From approve, the answer yes leads to rollback and the answer no leads to report. Rollback leads to report, and report leads to END. The graph is illustrative.STARTgather_evidencedecideapproveinterrupt()rollbackreportENDmonitoryesno
The LangGraph incident triage graph

For example, a LangGraph incident graph gathers evidence about a 5xx spike on the checkout service, asks Claude for a decision, and pauses at an approval node before any production rollback.

This LangGraph Python example builds the same agent as the Claude Agent SDK tutorial and the LangChain tutorial. The difference here is structural: evidence gathering always runs in a fixed order, the model only makes the decision, and the approval step is a node in the graph rather than a check hidden inside a tool.

Prerequisites

  • Python 3.11: LangGraph requires Python 3.10 or later, and the examples target version 3.11.
  • Tool module: The incident_tools.py file from Step 1 of the LangChain tutorial, which provides the three read-only tools.
  • Approval concepts: Familiarity with human-in-the-loop design and with structured output from language models.
  • Credentials: An Anthropic API key stored in the ANTHROPIC_API_KEY environment variable, never written into the code.

Setup

Create a virtual environment and install LangGraph with the Anthropic chat model integration at pinned versions.

Bash
python3.11 -m venv .venv
source .venv/bin/activate
pip install langgraph==1.0.0 langchain-anthropic==1.0.0
export ANTHROPIC_API_KEY="paste-your-key-here"

Step 1: Plan the Graph in Plain Python

Designing the routing without any library first makes every execution path testable and deterministic. The program below uses the same node names and routing rules as the LangGraph implementation, with a fixed rule substituting for the language model's decision.

Python
# graph_plan.py: the incident graph as plain Python, with no model and no LangGraph.
def gather_evidence(state):
    return {"evidence": "502 timeouts, 5xx at 12.0%, v2.4.1 deployed at 14:03"}

def decide(state):
    # Stand-in for the model: a deploy just before the errors points to a rollback.
    return {"action": "rollback", "version": "v2.4.1"}

def approve(state):
    return {"approved": state["human_answer"] == "yes"}

def rollback(state):
    return {"result": f"rolled back from {state['version']}"}

def report(state):
    return {"result": state.get("result", "no change made, report filed")}

NODES = {"gather_evidence": gather_evidence, "decide": decide,
         "approve": approve, "rollback": rollback, "report": report}
ROUTES = {
    "gather_evidence": lambda s: "decide",
    "decide": lambda s: "approve" if s["action"] == "rollback" else "report",
    "approve": lambda s: "rollback" if s["approved"] else "report",
    "rollback": lambda s: "report",
    "report": lambda s: "END",
}

def run(human_answer):
    state = {"service": "checkout", "human_answer": human_answer}
    node, path = "gather_evidence", []
    while node != "END":
        path.append(node)
        state.update(NODES[node](state))
        node = ROUTES[node](state)
    print(f"human answer {human_answer!r}: {' -> '.join(path)}")
    print(f"  result: {state['result']}")

run("yes")
run("no")

Output:

Example
human answer 'yes': gather_evidence -> decide -> approve -> rollback -> report
  result: rolled back from v2.4.1
human answer 'no': gather_evidence -> decide -> approve -> report
  result: no change made, report filed
  • Partial updates: Each node returns only the keys it modifies, and the runner merges them into the shared state, which mirrors the default LangGraph behaviour.
  • Guaranteed routing: A rejected approval proceeds directly to report, so the rollback node is structurally unreachable without an affirmative answer.
  • Testability: Both branches can be verified in continuous integration before any model, credential or network connection is involved.

Step 2: Define the State and the Nodes

The state is declared as a TypedDict. The decide node requests a validated Decision object through structured output, which constrains the model to two permitted actions. The approve node calls interrupt(), which suspends execution and subsequently returns the engineer's answer.

Python
# graph_nodes.py: state, model and node functions for the incident graph.
from typing import Literal, TypedDict

from langchain_anthropic import ChatAnthropic
from langgraph.types import interrupt
from pydantic import BaseModel

import incident_tools as it

class IncidentState(TypedDict, total=False):
    service: str
    evidence: dict
    cause: str
    action: str
    version: str
    approved: bool
    result: str

class Decision(BaseModel):
    cause: str
    action: Literal["rollback", "monitor"]
    version: str

decider = ChatAnthropic(model="claude-sonnet-5", temperature=0).with_structured_output(Decision)

def gather_evidence(state: IncidentState):
    s = state["service"]
    return {"evidence": {"logs": it.read_logs(s), "metrics": it.get_metrics(s),
                         "deploys": it.list_recent_deploys(s)}}

def decide(state: IncidentState):
    d = decider.invoke(
        "You triage production incidents. Name the likely cause, then choose "
        "rollback (with the version) or monitor (with an empty version).\n"
        f"Evidence: {state['evidence']}"
    )
    return {"cause": d.cause, "action": d.action, "version": d.version}

def approve(state: IncidentState):
    answer = interrupt({"question": f"Roll back {state['service']} from {state['version']}?",
                        "cause": state["cause"]})
    return {"approved": answer == "yes"}

def rollback(state: IncidentState):
    return {"result": f"{state['service']} rolled back from {state['version']}"}

def report(state: IncidentState):
    return {"result": state.get("result", f"No change made. Cause: {state['cause']}")}
  • Deterministic evidence: The gather_evidence node always calls all three tools, so the model cannot skip a data source before deciding.
  • Re-execution on resumption: LangGraph executes the approve node again from its beginning when the run resumes, so no side effect should precede interrupt().
  • Isolated action: The rollback operation occupies a dedicated node, which only the approved branch can reach.

Step 3: Build the LangGraph StateGraph with Conditional Edges

The StateGraph builder registers the nodes, unconditional edges connect the sequential steps, and LangGraph conditional edges, added with add_conditional_edges(), invoke a routing function that returns the name of the next node. Compiling the graph with a checkpointer is mandatory for interrupts, because the suspended state must be persisted somewhere.

Python
# incident_graph.py: wire the nodes into a compiled graph.
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import END, START, StateGraph

from graph_nodes import IncidentState, approve, decide, gather_evidence, report, rollback

def after_decide(state: IncidentState) -> str:
    return "approve" if state["action"] == "rollback" else "report"

def after_approve(state: IncidentState) -> str:
    return "rollback" if state["approved"] else "report"

builder = StateGraph(IncidentState)
for name, fn in [("gather_evidence", gather_evidence), ("decide", decide),
                 ("approve", approve), ("rollback", rollback), ("report", report)]:
    builder.add_node(name, fn)

builder.add_edge(START, "gather_evidence")
builder.add_edge("gather_evidence", "decide")
builder.add_conditional_edges("decide", after_decide, ["approve", "report"])
builder.add_conditional_edges("approve", after_approve, ["rollback", "report"])
builder.add_edge("rollback", "report")
builder.add_edge("report", END)

graph = builder.compile(checkpointer=InMemorySaver())

Step 4: Run, Pause and Resume

The first invoke() call stops at the interrupt and returns the pending request under the __interrupt__ key. The second call supplies Command(resume=answer) with the identical thread identifier, so the graph continues from the persisted checkpoint instead of starting again.

Python
# run_triage.py: start the graph, ask the engineer, then resume.
from langgraph.types import Command

from incident_graph import graph

config = {"configurable": {"thread_id": "incident-checkout-001"}}
first = graph.invoke({"service": "checkout"}, config)

if "__interrupt__" in first:
    request = first["__interrupt__"][0].value
    print(f"Cause: {request['cause']}")
    answer = input(f"{request['question']} [yes/no] ").strip().lower()
    final = graph.invoke(Command(resume=answer), config)
else:
    final = first
print(f"Result: {final['result']}")

Output

The LangGraph code was not executed for this page, so the following output is illustrative. The model's description of the cause varies between executions, while the path through the graph is determined entirely by the edges.

Example
Cause: Deploy v2.4.1 at 14:03 cut the payment client timeout from 5000 ms to
2000 ms; 502 timeouts began at 14:05 and the 5xx rate is 12.0% (baseline 0.2%).
Roll back checkout from v2.4.1? [yes/no] yes
Result: checkout rolled back from v2.4.1
  • Durable suspension: Between the two calls, the state exists only in the checkpointer, so a production deployment would require a database-backed implementation.
  • Predictable cost: The graph performs exactly one model call, whereas the agents in the other two tutorials determine their own number of calls.
  • Auditability: Every checkpoint records the evidence, the decision and the approval, which provides a complete history for the incident review.

Common Errors

  • ValueError: Checkpointer requires one or more of the following 'configurable' keys: The configuration lacks a thread_id, so pass the identical configuration dictionary to both invocations.
  • KeyError: '__interrupt__': The model chose monitor, so the graph finished without pausing; check for the key before reading it, as Step 4 does.
  • ValueError saying a name is already being used as a state key: A node shares its name with a state field, which the builder rejects during registration, so rename the node.
  • InvalidUpdateError: Two parallel nodes wrote the same key during one step, so annotate the key with a reducer function or execute the nodes sequentially.
  • GraphRecursionError: A cycle exceeded the recursion limit without reaching END, so confirm that every loop contains a reachable termination condition.
  • The run starts again from the beginning: The resumption call used a different thread identifier, so the checkpointer found no suspended execution and treated the input as a new run.

Next Steps

  • Multi-agent graphs: Split triage and remediation into separate subgraphs, as described in multi-agent systems.
  • Role-based alternative: Build the same workflow with collaborating role agents in the CrewAI tutorial.
  • Framework selection: Compare control, persistence and ergonomics in LangGraph vs CrewAI vs AutoGen.
  • Quality measurement: Replay recorded incidents through the graph and score the decisions with AI agent evaluation.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In the LangGraph triage agent, what pauses the run before a rollback?

Frequently Asked Questions

What is LangGraph used for?

LangGraph is used to build agents and workflows whose control flow must be explicit, such as branching, loops, retries and human approval. The state is saved at every step, so a run can pause, resume or be inspected later.

Is LangGraph part of LangChain?

LangGraph is a separate library from the same team. It works with LangChain models and tools but does not require them, and the LangChain agent constructor is itself built on LangGraph.

What is a checkpointer in LangGraph?

A checkpointer saves the graph state after each step under a thread ID. It is required for interrupts, because the paused run must be stored until a person responds, and production systems use a database-backed checkpointer instead of memory.

When should LangGraph be chosen over a prebuilt agent?

A prebuilt agent suits tasks where the model can decide every step. LangGraph suits tasks where some steps must always happen, happen in a fixed order, or wait for a person. LangGraph human-in-the-loop steps, such as an approval before a production change, combine an interrupt with a checkpointer.