MCP Architecture: Host, Client, Server
MCP architecture is the client-server design of the Model Context Protocol, built from three roles: a host, one or more clients and one or more servers. The host runs the AI application, each client maintains a dedicated connection to one server, and each server exposes tools, resources and prompts from an external system.
- Host: The application that contains the language model, manages the conversation and decides which servers are connected.
- Client: A protocol component created by the host, holding a stateful one-to-one session with exactly one server.
- Server: An independent program that wraps an external system and answers protocol requests about its capabilities.
- Data layer: The JSON-RPC 2.0 messages for lifecycle management, discovery, tool invocation and notifications.
- Transport layer: The channel that carries those messages, either stdio for local processes or streamable HTTP for remote services.
For example, an AI assistant acting as the host creates three clients, connected respectively to an issue tracker server, a PostgreSQL server and a logs server, and routes each tool request to the correct client.
Key Characteristics of MCP Architecture
- Isolation: Each server runs separately and cannot read another server's data or the complete conversation history, which limits the damage from prompt injection.
- One client per server: The host creates a separate client for every server, so sessions, failures and permissions remain independent.
- Host-controlled context: The host decides which tool results and resources reach the model, which keeps the security boundary in the application.
- Capability negotiation: During initialisation, client and server declare the features they support, such as tools, resources or prompts.
- Transport independence: The same JSON-RPC messages work over stdio or streamable HTTP, so a server can move from a laptop to a remote service.
- Bidirectional messaging: Servers can also send notifications and certain requests back to the client, for example when their tool list changes.
How MCP Architecture Works
- Start-up: The host reads its configuration and creates one client for each registered server.
- Transport connection: A client either launches a local server as a subprocess and communicates over stdin and stdout, or connects to a remote server's HTTP endpoint.
- Initialisation: The client and server exchange protocol versions and capabilities, then the client confirms with an
initializednotification. - Aggregation: The host collects the tool lists from every client and presents one combined catalogue to the model.
- Routing: When the model requests a tool, the host identifies the owning client and forwards a
tools/callrequest. - Return path: The server's result travels back through the same client, and the host appends it to the model's context.
- Shutdown: The host closes each session when the conversation ends, which terminates local server processes.
Example: Tracing a Request Through Host, Client and Server
The program below is a teaching model that traces one tool request through the three roles, using two in-memory servers.
# A trace of one request through host, client and server (a teaching model).
# The initialize handshake is left out to keep the trace short.
class Server:
def __init__(self, name, tools):
self.name, self.tools = name, tools
def handle(self, msg):
if msg["method"] == "tools/list":
return {"tools": list(self.tools)}
if msg["method"] == "tools/call":
fn = self.tools[msg["params"]["name"]]
return {"content": [{"type": "text", "text": fn(**msg["params"]["arguments"])}]}
class Client:
# One client per server: it holds a single 1:1 connection.
def __init__(self, server):
self.server, self.next_id = server, 1
def request(self, method, params=None):
msg = {"jsonrpc": "2.0", "id": self.next_id, "method": method, "params": params or {}}
self.next_id += 1
print(f" client -> {self.server.name}: {method} (id {msg['id']})")
result = self.server.handle(msg)
print(f" {self.server.name} -> client: result for id {msg['id']}")
return result
class Host:
def __init__(self, servers):
self.clients = [Client(s) for s in servers]
self.routes = {} # tool name -> client that can call it
def connect(self):
for c in self.clients:
for tool in c.request("tools/list")["tools"]:
self.routes[tool] = c
print("host: tools available to the model:", sorted(self.routes))
def run_tool(self, name, arguments):
print(f"host: model asked for {name} {arguments}")
result = self.routes[name].request("tools/call", {"name": name, "arguments": arguments})
return result["content"][0]["text"]
issues = Server("issues", {"get_issue": lambda issue_id: f"{issue_id}: checkout 500 on empty cart (open)"})
logs = Server("logs", {"recent_errors": lambda service: f"{service}: 42 HTTP 500 errors in the last 15 minutes"})
host = Host([issues, logs])
host.connect()
print("host: text returned to model:", host.run_tool("recent_errors", {"service": "checkout"})) client -> issues: tools/list (id 1)
issues -> client: result for id 1
client -> logs: tools/list (id 1)
logs -> client: result for id 1
host: tools available to the model: ['get_issue', 'recent_errors']
host: model asked for recent_errors {'service': 'checkout'}
client -> logs: tools/call (id 2)
logs -> client: result for id 2
host: text returned to model: checkout: 42 HTTP 500 errors in the last 15 minutes- Separate sessions: Each client numbers its own requests, which is why both
tools/listcalls use id 1. - Routing table: The host maps every tool name to the client that discovered it, so
recent_errorsreaches only the logs server. - Single gateway: The model never contacts a server directly; every request and result passes through the host.
Applications of MCP Architecture
- Multi-tool assistants: Combining an issue tracker, a relational database and a logs service within one debugging session.
- Local developer tools: Running file system or version control servers as subprocesses on a developer workstation.
- Remote enterprise services: Hosting shared servers over streamable HTTP behind an organisation's authentication.
- Agent platforms: Supplying tools to AI agent architecture designs without custom adapters for every system.
- Sandboxed execution: Running untrusted servers in containers while the host enforces access policies.
Advantages
- Fault isolation: A crashed or slow server affects only its own client, not the whole application.
- Least privilege: Each server receives only the credentials it needs, such as a read-only database account.
- Incremental adoption: Teams can add or remove servers through configuration without modifying the host application.
- Clear responsibilities: Hosts manage users and models, while servers manage integrations with specific systems.
Limitations
- Process overhead: Every local server is a separate process, which increases memory usage on a workstation.
- Name collisions: Two servers may expose tools with identical names, so hosts need a disambiguation strategy.
- Distributed debugging: Failures can occur in the host, the client, the transport or the server, which complicates troubleshooting.
- Authorisation complexity: Remote servers require authentication and token management, which the host and server must implement correctly.
The difference between this design and a direct integration is explained in MCP vs API vs function calling. Background on the protocol itself appears in what is MCP, the model-side mechanism is tool calling, and a working server is implemented in build an MCP server in Python.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. In MCP architecture, how many clients does a host create to connect to three servers?
Frequently Asked Questions
MCP host vs client: what is the difference?
The host is the whole AI application that the user interacts with. A client is a component inside the host that manages one connection to one server, so a host with three servers contains three clients.
Can one MCP client connect to several servers?
No. Each client holds a one-to-one session with a single server. A host that needs several servers creates one client for each of them.
When should MCP stdio be used instead of streamable HTTP?
Stdio suits a server that runs on the same machine as the host, because the host can launch it as a subprocess. Streamable HTTP suits a shared server that many users reach over a network.
Does the language model talk to MCP servers directly?
No. The model only requests a tool, and the host forwards that request through the correct client. The host also decides which results are added to the model's context.
Related Articles
- What is MCP (Model Context Protocol)Learn what the Model Context Protocol (MCP) is: hosts, clients and servers, tools and resources, JSON-RPC messages, and a Python teaching model of MCP.
- MCP vs API vs Function CallingMCP vs API vs function calling explained: the layer each one works at, a comparison table, when to use each, and one issue tracker search done three ways.
- Build an MCP Server in PythonBuild an MCP server in Python: a standard-library JSON-RPC teaching server, the same server with the MCP SDK, host configuration and common errors fixed.
- AI Agent ArchitectureAI agent architecture explained: the model, tools, memory, control loop and guardrails, how they interact, and a traced Python skeleton for a DevOps agent.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- A2A Protocol vs MCPA2A protocol vs MCP: agent-to-agent delegation versus agent-to-tool calls, a comparison table, when to use each, and an incident example that uses both.