…
Skip to content
Topics
On this page

MCP Architecture: Host, Client, Server

MCP architecture is the client-server design of the Model Context Protocol, built from three roles: a host, one or more clients and one or more servers. The host runs the AI application, each client maintains a dedicated connection to one server, and each server exposes tools, resources and prompts from an external system.

  • Host: The application that contains the language model, manages the conversation and decides which servers are connected.
  • Client: A protocol component created by the host, holding a stateful one-to-one session with exactly one server.
  • Server: An independent program that wraps an external system and answers protocol requests about its capabilities.
  • Data layer: The JSON-RPC 2.0 messages for lifecycle management, discovery, tool invocation and notifications.
  • Transport layer: The channel that carries those messages, either stdio for local processes or streamable HTTP for remote services.
MCP host, clients and serversA host box on the left contains the LLM and three MCP clients. Each client has its own two-way connection to one MCP server on the right: an issue tracker server and a PostgreSQL server over stdio, and a logs server over streamable HTTP. The transport choices are illustrative.Host: AI assistantLLMClient 1stdioIssue tracker serverClient 2stdioPostgreSQL serverClient 3HTTPLogs server
MCP host, clients and servers

For example, an AI assistant acting as the host creates three clients, connected respectively to an issue tracker server, a PostgreSQL server and a logs server, and routes each tool request to the correct client.

Key Characteristics of MCP Architecture

  • Isolation: Each server runs separately and cannot read another server's data or the complete conversation history, which limits the damage from prompt injection.
  • One client per server: The host creates a separate client for every server, so sessions, failures and permissions remain independent.
  • Host-controlled context: The host decides which tool results and resources reach the model, which keeps the security boundary in the application.
  • Capability negotiation: During initialisation, client and server declare the features they support, such as tools, resources or prompts.
  • Transport independence: The same JSON-RPC messages work over stdio or streamable HTTP, so a server can move from a laptop to a remote service.
  • Bidirectional messaging: Servers can also send notifications and certain requests back to the client, for example when their tool list changes.

How MCP Architecture Works

  1. Start-up: The host reads its configuration and creates one client for each registered server.
  2. Transport connection: A client either launches a local server as a subprocess and communicates over stdin and stdout, or connects to a remote server's HTTP endpoint.
  3. Initialisation: The client and server exchange protocol versions and capabilities, then the client confirms with an initialized notification.
  4. Aggregation: The host collects the tool lists from every client and presents one combined catalogue to the model.
  5. Routing: When the model requests a tool, the host identifies the owning client and forwards a tools/call request.
  6. Return path: The server's result travels back through the same client, and the host appends it to the model's context.
  7. Shutdown: The host closes each session when the conversation ends, which terminates local server processes.

Example: Tracing a Request Through Host, Client and Server

The program below is a teaching model that traces one tool request through the three roles, using two in-memory servers.

Python
# A trace of one request through host, client and server (a teaching model).
# The initialize handshake is left out to keep the trace short.
class Server:
    def __init__(self, name, tools):
        self.name, self.tools = name, tools

    def handle(self, msg):
        if msg["method"] == "tools/list":
            return {"tools": list(self.tools)}
        if msg["method"] == "tools/call":
            fn = self.tools[msg["params"]["name"]]
            return {"content": [{"type": "text", "text": fn(**msg["params"]["arguments"])}]}

class Client:
    # One client per server: it holds a single 1:1 connection.
    def __init__(self, server):
        self.server, self.next_id = server, 1

    def request(self, method, params=None):
        msg = {"jsonrpc": "2.0", "id": self.next_id, "method": method, "params": params or {}}
        self.next_id += 1
        print(f"  client -> {self.server.name}: {method} (id {msg['id']})")
        result = self.server.handle(msg)
        print(f"  {self.server.name} -> client: result for id {msg['id']}")
        return result

class Host:
    def __init__(self, servers):
        self.clients = [Client(s) for s in servers]
        self.routes = {}  # tool name -> client that can call it

    def connect(self):
        for c in self.clients:
            for tool in c.request("tools/list")["tools"]:
                self.routes[tool] = c
        print("host: tools available to the model:", sorted(self.routes))

    def run_tool(self, name, arguments):
        print(f"host: model asked for {name} {arguments}")
        result = self.routes[name].request("tools/call", {"name": name, "arguments": arguments})
        return result["content"][0]["text"]

issues = Server("issues", {"get_issue": lambda issue_id: f"{issue_id}: checkout 500 on empty cart (open)"})
logs = Server("logs", {"recent_errors": lambda service: f"{service}: 42 HTTP 500 errors in the last 15 minutes"})

host = Host([issues, logs])
host.connect()
print("host: text returned to model:", host.run_tool("recent_errors", {"service": "checkout"}))
Output
  client -> issues: tools/list (id 1)
  issues -> client: result for id 1
  client -> logs: tools/list (id 1)
  logs -> client: result for id 1
host: tools available to the model: ['get_issue', 'recent_errors']
host: model asked for recent_errors {'service': 'checkout'}
  client -> logs: tools/call (id 2)
  logs -> client: result for id 2
host: text returned to model: checkout: 42 HTTP 500 errors in the last 15 minutes
  • Separate sessions: Each client numbers its own requests, which is why both tools/list calls use id 1.
  • Routing table: The host maps every tool name to the client that discovered it, so recent_errors reaches only the logs server.
  • Single gateway: The model never contacts a server directly; every request and result passes through the host.

Applications of MCP Architecture

  • Multi-tool assistants: Combining an issue tracker, a relational database and a logs service within one debugging session.
  • Local developer tools: Running file system or version control servers as subprocesses on a developer workstation.
  • Remote enterprise services: Hosting shared servers over streamable HTTP behind an organisation's authentication.
  • Agent platforms: Supplying tools to AI agent architecture designs without custom adapters for every system.
  • Sandboxed execution: Running untrusted servers in containers while the host enforces access policies.

Advantages

  • Fault isolation: A crashed or slow server affects only its own client, not the whole application.
  • Least privilege: Each server receives only the credentials it needs, such as a read-only database account.
  • Incremental adoption: Teams can add or remove servers through configuration without modifying the host application.
  • Clear responsibilities: Hosts manage users and models, while servers manage integrations with specific systems.

Limitations

  • Process overhead: Every local server is a separate process, which increases memory usage on a workstation.
  • Name collisions: Two servers may expose tools with identical names, so hosts need a disambiguation strategy.
  • Distributed debugging: Failures can occur in the host, the client, the transport or the server, which complicates troubleshooting.
  • Authorisation complexity: Remote servers require authentication and token management, which the host and server must implement correctly.

The difference between this design and a direct integration is explained in MCP vs API vs function calling. Background on the protocol itself appears in what is MCP, the model-side mechanism is tool calling, and a working server is implemented in build an MCP server in Python.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. In MCP architecture, how many clients does a host create to connect to three servers?

Frequently Asked Questions

MCP host vs client: what is the difference?

The host is the whole AI application that the user interacts with. A client is a component inside the host that manages one connection to one server, so a host with three servers contains three clients.

Can one MCP client connect to several servers?

No. Each client holds a one-to-one session with a single server. A host that needs several servers creates one client for each of them.

When should MCP stdio be used instead of streamable HTTP?

Stdio suits a server that runs on the same machine as the host, because the host can launch it as a subprocess. Streamable HTTP suits a shared server that many users reach over a network.

Does the language model talk to MCP servers directly?

No. The model only requests a tool, and the host forwards that request through the correct client. The host also decides which results are added to the model's context.