Claude Agent SDK Tutorial
The Claude Agent SDK is a Python and TypeScript library for building autonomous agents on the same runtime that powers Claude Code. It executes the agent loop, calls tools, manages context and enforces permissions, so application code only defines the tools, the instructions and the safety policy.
- Orchestration: The SDK sends the goal to Claude, executes each requested tool and returns the observation until the task is complete.
- Extensibility: Ordinary Python functions become callable tools through a decorator and an in-process MCP server.
- Authorisation: A callback named
can_use_tooldecides whether each unlisted tool invocation is permitted to execute. - Observability: Every run produces typed, streaming messages, so the application can display each tool invocation and each intermediate answer.
- Built-in capabilities: File, shell and search tools from Claude Code are available by default, and the application can restrict them.
For example, an on-call engineer asks the agent to investigate a 5xx spike on the checkout service, and the agent examines logs, metrics and deployments, then proposes a rollback that requires human approval.
The incident triage agent below is a complete Claude Agent SDK example, and an identical agent is implemented in the LangChain tutorial and the LangGraph tutorial, which makes the three frameworks directly comparable. The agent has three read-only tools (read_logs, get_metrics and list_recent_deploys) and one state-changing tool, rollback_deploy, which must never execute without explicit approval from a person.
Prerequisites
- Python 3.11: The SDK requires Python 3.10 or later, and the examples were written for version 3.11.
- Agent fundamentals: Familiarity with AI agents and tool calling, including how a model requests a function with arguments.
- Protocol fundamentals: A general understanding of MCP, because the SDK exposes custom tools through an MCP server.
- Credentials: An Anthropic API key stored in the
ANTHROPIC_API_KEYenvironment variable, never written into the source code.
Setup
Create an isolated virtual environment and install the pinned version of the Claude Agent SDK Python package. The Python package communicates with the Claude Code runtime, which recent releases include automatically.
python3.11 -m venv .venv
source .venv/bin/activate
pip install claude-agent-sdk==0.1.0
export ANTHROPIC_API_KEY="paste-your-key-here"Step 1: Write the Three Incident Tools
The tools and the permission policy are written as ordinary Python first, so their behaviour can be tested deterministically without any model or network access. Save the following module as incident_tools.py; the other files import it.
# incident_tools.py: canned incident data, the three tools and the permission rule.
import json
LOGS = {
"checkout": [
"14:05:12 ERROR POST /api/pay 502 payment client timeout after 2000 ms",
"14:05:13 ERROR POST /api/pay 502 payment client timeout after 2000 ms",
"14:06:40 WARN retry budget exhausted for payments-api",
],
}
METRICS = {"checkout": {"5xx_rate": "12.0%", "baseline_5xx_rate": "0.2%", "p95_latency_ms": 2140}}
DEPLOYS = {
"checkout": [
{"version": "v2.4.1", "at": "14:03", "change": "payment client timeout 5000 ms to 2000 ms"},
],
}
def read_logs(service):
return "\n".join(LOGS.get(service, ["no log lines found"]))
def get_metrics(service):
return json.dumps(METRICS.get(service, {}))
def list_recent_deploys(service):
return json.dumps(DEPLOYS.get(service, []))
READ_ONLY = {"read_logs", "get_metrics", "list_recent_deploys"}
NEEDS_APPROVAL = {"rollback_deploy"}
def permission(tool, approver_says=None):
# Deny by default: read-only tools run, a rollback needs a human "yes".
if tool in READ_ONLY:
return "allow"
if tool in NEEDS_APPROVAL:
return "allow" if approver_says == "yes" else "deny"
return "deny"
if __name__ == "__main__":
for fn in (read_logs, get_metrics, list_recent_deploys):
print(f"{fn.__name__}('checkout'):\n{fn('checkout')}\n")
checks = [("get_metrics", None), ("rollback_deploy", None), ("rollback_deploy", "yes"), ("Bash", None)]
for tool, answer in checks:
print(f"{tool:<16} approver={str(answer):<4} -> {permission(tool, answer)}")Output:
read_logs('checkout'):
14:05:12 ERROR POST /api/pay 502 payment client timeout after 2000 ms
14:05:13 ERROR POST /api/pay 502 payment client timeout after 2000 ms
14:06:40 WARN retry budget exhausted for payments-api
get_metrics('checkout'):
{"5xx_rate": "12.0%", "baseline_5xx_rate": "0.2%", "p95_latency_ms": 2140}
list_recent_deploys('checkout'):
[{"version": "v2.4.1", "at": "14:03", "change": "payment client timeout 5000 ms to 2000 ms"}]
get_metrics approver=None -> allow
rollback_deploy approver=None -> deny
rollback_deploy approver=yes -> allow
Bash approver=None -> deny- Canned evidence: The logs contain payment client timeouts that begin shortly after deployment v2.4.1 reduced the payment timeout, which is the correlation the agent should identify.
- Default denial: Any tool outside the two named sets is refused, including built-in tools such as
Bash, so an unexpected capability can never execute silently. - Separation of concerns: The policy function contains no SDK code, so the same rule can be reused and unit tested in any framework.
Step 2: Register the Tools on an MCP Server
The @tool decorator accepts a name, a description and an input schema. Each asynchronous handler receives its arguments as a dictionary and returns MCP content blocks. The create_sdk_mcp_server() function then serves all four tools inside the Python process, so no separate server or network port is necessary.
# agent_tools.py: the four tools as Claude Agent SDK tools.
from claude_agent_sdk import create_sdk_mcp_server, tool
import incident_tools as it
def as_text(result):
return {"content": [{"type": "text", "text": result}]}
@tool("read_logs", "Read recent error log lines for a service", {"service": str})
async def read_logs(args):
return as_text(it.read_logs(args["service"]))
@tool("get_metrics", "Get the current 5xx rate and latency for a service", {"service": str})
async def get_metrics(args):
return as_text(it.get_metrics(args["service"]))
@tool("list_recent_deploys", "List deploys of a service in the last 24 hours", {"service": str})
async def list_recent_deploys(args):
return as_text(it.list_recent_deploys(args["service"]))
@tool("rollback_deploy", "Roll a production service back from a version", {"service": str, "version": str})
async def rollback_deploy(args):
return as_text(f"{args['service']} rolled back from {args['version']}")
server = create_sdk_mcp_server(
name="incident",
version="1.0.0",
tools=[read_logs, get_metrics, list_recent_deploys, rollback_deploy],
)Step 3: Add the Permission Callback for Rollbacks
The SDK names each server tool with the prefix mcp__incident__ followed by the function name. The three read-only tools appear in allowed_tools, so they execute immediately. The rollback tool is deliberately excluded, so every rollback request invokes can_use_tool, which asks the on-call engineer for a decision in the terminal.
# agent_options.py: model, instructions, tools and the approval rule.
import asyncio
from claude_agent_sdk import ClaudeAgentOptions
from claude_agent_sdk.types import PermissionResultAllow, PermissionResultDeny
import incident_tools as it
from agent_tools import server
PREFIX = "mcp__incident__"
SYSTEM = (
"You are an on-call triage agent. Use the incident tools to find the likely cause "
"of an error spike, then propose one action. Call rollback_deploy only when a "
"recent deploy explains the errors. Never guess data you have not read."
)
async def can_use_tool(tool_name, input_data, context):
name = tool_name.removeprefix(PREFIX)
answer = None
if name in it.NEEDS_APPROVAL:
prompt = f"Approve {name} {input_data}? [yes/no] "
answer = (await asyncio.to_thread(input, prompt)).strip().lower()
if it.permission(name, answer) == "allow":
return PermissionResultAllow()
return PermissionResultDeny(message=f"{name} was not approved by the on-call engineer")
options = ClaudeAgentOptions(
model="claude-sonnet-5",
system_prompt=SYSTEM,
mcp_servers={"incident": server},
allowed_tools=[PREFIX + name for name in sorted(it.READ_ONLY)],
can_use_tool=can_use_tool,
max_turns=10,
)- Non-blocking input:
asyncio.to_thread()preventsinput()from blocking the event loop that delivers streaming messages. - Denial feedback: The message inside
PermissionResultDenyis returned to Claude, so the agent can report the rejection rather than repeating the request. - Built-in restrictions: The callback denies built-in tools by default, and the
disallowed_toolsoption can also remove them from the available tool list. - Turn limit: The
max_turnssetting stops a looping agent before it consumes an unreasonable number of tokens.
Step 4: Run the Claude Agent SDK Agent
The ClaudeSDKClient class opens an interactive session in streaming mode, which the permission callback requires. The program sends the goal, then iterates over the response messages and prints every tool invocation, every text block and the final number of turns.
# triage_agent.py: run the incident triage agent.
import asyncio
from claude_agent_sdk import AssistantMessage, ClaudeSDKClient, ResultMessage, TextBlock, ToolUseBlock
from agent_options import options
GOAL = "The checkout service has a 5xx spike. Find the likely cause and propose an action."
async def main():
async with ClaudeSDKClient(options=options) as client:
await client.query(GOAL)
async for message in client.receive_response():
if isinstance(message, AssistantMessage):
for block in message.content:
if isinstance(block, ToolUseBlock):
print(f"Tool call: {block.name} {block.input}")
elif isinstance(block, TextBlock):
print(f"Claude: {block.text}")
elif isinstance(message, ResultMessage):
print(f"Finished after {message.num_turns} turns")
asyncio.run(main())Output
The SDK code above was not executed for this page, so the following output is illustrative. The wording, the order of tool invocations and the number of turns vary between runs, because the model decides each step.
Tool call: mcp__incident__get_metrics {'service': 'checkout'}
Tool call: mcp__incident__read_logs {'service': 'checkout'}
Tool call: mcp__incident__list_recent_deploys {'service': 'checkout'}
Claude: The 5xx rate is 12.0% against a 0.2% baseline. The errors are 502 payment
client timeouts that began at 14:05, two minutes after v2.4.1 cut the payment
timeout from 5000 ms to 2000 ms, while p95 latency is 2140 ms.
Tool call: mcp__incident__rollback_deploy {'service': 'checkout', 'version': 'v2.4.1'}
Approve rollback_deploy {'service': 'checkout', 'version': 'v2.4.1'}? [yes/no] no
Claude: The rollback was not approved. Recommended action: restore the 5000 ms
timeout in a hotfix or approve the rollback of v2.4.1.
Finished after 6 turns- Evidence first: The agent reads all three data sources before it names a cause, because the system prompt forbids guessing data it has not read.
- Approval enforced in code: The rollback reached the callback and stopped there, so the model could not bypass the rule by rewording its request.
- Useful fallback: After the denial, the agent still returns a recommendation that the engineer can act on.
Common Errors
CLINotFoundError: The SDK cannot locate the Claude Code runtime, so reinstall the package inside the active environment or install the command-line interface that the message identifies.ProcessError: The runtime exited with a non-zero status, frequently because the API key variable is missing or the model identifier is incorrect, so inspect the exit code and the error stream.CLIConnectionError: The client could not start or contact the runtime, so confirm that the asynchronous context manager surrounds everyquery()call.ValueErrorabout streaming mode:can_use_toolwas passed to the one-shotquery()function with a string prompt, so useClaudeSDKClientas in Step 4.- Every read tool requests approval: The names in
allowed_toolsomit themcp__incident__prefix, so no invocation matches the list and every call reaches the permission callback.
Next Steps
- Framework comparison: Rebuild the identical agent with decorated tools and a chat model integration in the LangChain tutorial.
- Explicit control flow: Represent the approval as a graph interruption with persistent state in the LangGraph tutorial.
- Framework selection: Evaluate the architectural trade-offs described in LangGraph vs CrewAI vs AutoGen.
- Security and quality: Protect tool output against prompt injection and measure agent behaviour with AI agent evaluation.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In the triage agent, why is rollback_deploy left out of allowed_tools?
Frequently Asked Questions
What is the difference between the Claude Agent SDK and the Anthropic SDK?
The Anthropic SDK sends single requests to the Messages API, and the application writes its own tool loop. The Claude Agent SDK runs the complete agent loop, including tool execution, context handling and permission checks, on the same runtime as Claude Code.
Does the Claude Agent SDK need Claude Code installed?
The SDK drives the Claude Code runtime as a subprocess. Recent releases of the Python package include that runtime, while older releases expected a separate installation, so the package documentation for the pinned version is the reference.
How do you add Claude Agent SDK custom tools in Python?
A Python function decorated with the tool decorator is registered on an in-process MCP server, and the agent calls it like any other tool. No separate server process or network port is required.
How do Claude Agent SDK permissions stop an agent from running a dangerous tool?
Tools listed in allowed_tools run without a check. Every other tool call goes to the can_use_tool callback, which can allow it, deny it with a message, or ask a person first, as the rollback step in the example does.
Related Articles
- LangChain TutorialLangChain tutorial: define tools with @tool, connect Claude through langchain-anthropic, build a tool-calling agent and a chain for incident triage.
- LangGraph TutorialLangGraph tutorial: build an incident triage agent as a StateGraph with nodes, conditional edges and a human approval interrupt before a rollback.
- Claude Code TutorialClaude Code tutorial: install it, start a session in a repository, investigate a 5xx incident, and use CLAUDE.md, permission modes, slash commands and MCP.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- Human-in-the-Loop in Agentic AIHuman-in-the-loop in agentic AI: risk-based approval gates, escalation and audit logs, with a Python approval queue for an incident rollback and limits.
- What is MCP (Model Context Protocol)Learn what the Model Context Protocol (MCP) is: hosts, clients and servers, tools and resources, JSON-RPC messages, and a Python teaching model of MCP.