Model Context Protocol (MCP): Complete Guide to Building Tool-Using AI Agents

Complete MCP implementation guide for building tool-using AI agents with Claude, covering server architecture, transport layers, and production patterns.

#claude#mcp#model-context-protocol#ai-agents#integration
Cover image for the article: Model Context Protocol (MCP): Complete Guide to Building Tool-Using AI Agents

Model Context Protocol (MCP) is the open standard that turns LLMs from text generators into tool-using agents. Before MCP, every AI tool integration was a custom implementation — bespoke function definitions, proprietary schemas, one-off connection handlers. MCP standardizes how AI models discover, invoke, and receive results from external tools, creating an ecosystem where tools are portable across models and platforms. This guide covers everything you need to build production MCP servers, from protocol architecture to deployment patterns, based on shipping six MCP servers that handle 2.3 million tool invocations per month.

What Is the Model Context Protocol and Why Does It Matter?

MCP defines a client-server protocol where AI applications (clients) connect to tool providers (servers). The server exposes capabilities — tools, resources, and prompts — through a standardized JSON-RPC interface. The client discovers these capabilities, presents them to the model, and routes tool invocations back to the server.

The protocol solves three problems that plagued pre-MCP tool integrations:

ProblemPre-MCP SolutionMCP Solution
Tool discoveryHardcoded tool lists in promptsDynamic capability discovery
Schema definitionPlatform-specific JSON schemasStandardized tool schemas
TransportCustom HTTP endpointsStdio, SSE, or HTTP with standard lifecycle
AuthenticationPer-integration authTransport-layer auth with standard patterns
VersioningBreaking changes on updateProtocol version negotiation
Context provisionPrompt engineeringResources and resource templates

MCP Architecture: Clients, Servers, and Transport Layers

The MCP architecture has three components:

Host: The AI application that the user interacts with (Claude Desktop, Kiro, an IDE plugin). The host manages MCP client instances.

Client: A protocol client within the host that maintains a 1:1 connection with an MCP server. Handles capability negotiation, message routing, and lifecycle management.

Server: A lightweight process that exposes tools, resources, and prompts. Servers are stateless between invocations and can be written in any language.

// MCP Server implementation in TypeScript
import { Server } from '@modelcontextprotocol/sdk/server/index.js';
import { StdioServerTransport } from '@modelcontextprotocol/sdk/server/stdio.js';
import {
  CallToolRequestSchema,
  ListToolsRequestSchema,
  ListResourcesRequestSchema,
  ReadResourceRequestSchema,
} from '@modelcontextprotocol/sdk/types.js';

const server = new Server(
  { name: 'production-database-tools', version: '1.4.0' },
  { capabilities: { tools: {}, resources: {} } }
);

// Register tool definitions — discovered by clients automatically
server.setRequestHandler(ListToolsRequestSchema, async () => ({
  tools: [
    {
      name: 'query_production_metrics',
      description: 'Execute read-only queries against the production metrics database',
      inputSchema: {
        type: 'object',
        properties: {
          query: {
            type: 'string',
            description: 'SQL query (SELECT only, max 1000 rows)',
          },
          timeRange: {
            type: 'string',
            enum: ['1h', '6h', '24h', '7d', '30d'],
            description: 'Time range filter applied to all queries',
          },
        },
        required: ['query', 'timeRange'],
      },
    },
    {
      name: 'get_service_health',
      description: 'Check health status of a specific production service',
      inputSchema: {
        type: 'object',
        properties: {
          serviceName: { type: 'string', description: 'Service identifier' },
        },
        required: ['serviceName'],
      },
    },
  ],
}));

// Handle tool invocations
server.setRequestHandler(CallToolRequestSchema, async (request) => {
  const { name, arguments: args } = request.params;

  switch (name) {
    case 'query_production_metrics':
      return await handleMetricsQuery(args.query, args.timeRange);
    case 'get_service_health':
      return await handleHealthCheck(args.serviceName);
    default:
      throw new Error(`Unknown tool: ${name}`);
  }
});

// Start server with stdio transport
const transport = new StdioServerTransport();
await server.connect(transport);

How Do MCP Transport Layers Work?

MCP supports three transport mechanisms, each suited to different deployment models:

Stdio Transport

The server runs as a child process of the client. Communication happens over standard input/output streams. This is the simplest model and the default for local tools.

When to use: Local development tools, file system access, CLI wrappers. Latency: Sub-millisecond (no network overhead). Limitation: Cannot share a server across multiple clients.

SSE (Server-Sent Events) Transport

The server runs as an HTTP service. The client connects via SSE for server-to-client messages and POST requests for client-to-server messages.

When to use: Shared servers, remote tools, multi-tenant scenarios. Latency: 5-50ms depending on network. Limitation: Requires HTTP infrastructure and session management.

Streamable HTTP Transport

The newest transport option, supporting bidirectional streaming over HTTP. Designed for high-throughput scenarios where both sides need to send data continuously.

When to use: Real-time data feeds, large result sets, streaming tool outputs.

TransportLatencyDeploymentMulti-ClientStreaming
Stdio<1msLocal processNoNo
SSE5-50msHTTP serverYesServer→Client
Streamable HTTP5-50msHTTP serverYesBidirectional

Building a Production MCP Server: Step-by-Step

Here is a complete example of building an MCP server that provides infrastructure observability tools:

# infrastructure_mcp_server.py
import asyncio
import json
from mcp.server import Server
from mcp.server.stdio import stdio_server
from mcp.types import Tool, TextContent, Resource

app = Server("infrastructure-observability")

@app.list_tools()
async def list_tools() -> list[Tool]:
    return [
        Tool(
            name="get_pod_status",
            description="Get Kubernetes pod status for a namespace",
            inputSchema={
                "type": "object",
                "properties": {
                    "namespace": {"type": "string"},
                    "labelSelector": {"type": "string", "description": "Optional k8s label selector"},
                },
                "required": ["namespace"],
            },
        ),
        Tool(
            name="get_error_rate",
            description="Get error rate (5xx) for a service over a time window",
            inputSchema={
                "type": "object",
                "properties": {
                    "service": {"type": "string"},
                    "window": {"type": "string", "enum": ["5m", "15m", "1h", "6h"]},
                },
                "required": ["service", "window"],
            },
        ),
        Tool(
            name="get_resource_utilization",
            description="Get CPU and memory utilization for a service",
            inputSchema={
                "type": "object",
                "properties": {
                    "service": {"type": "string"},
                    "metric": {"type": "string", "enum": ["cpu", "memory", "both"]},
                },
                "required": ["service"],
            },
        ),
    ]

@app.call_tool()
async def call_tool(name: str, arguments: dict) -> list[TextContent]:
    match name:
        case "get_pod_status":
            result = await fetch_pod_status(
                arguments["namespace"],
                arguments.get("labelSelector"),
            )
        case "get_error_rate":
            result = await fetch_error_rate(
                arguments["service"],
                arguments["window"],
            )
        case "get_resource_utilization":
            result = await fetch_resource_utilization(
                arguments["service"],
                arguments.get("metric", "both"),
            )
        case _:
            return [TextContent(type="text", text=f"Unknown tool: {name}")]

    return [TextContent(type="text", text=json.dumps(result, indent=2))]

@app.list_resources()
async def list_resources() -> list[Resource]:
    return [
        Resource(
            uri="infra://runbooks/incident-response",
            name="Incident Response Runbook",
            description="Standard operating procedures for production incidents",
            mimeType="text/markdown",
        ),
    ]

async def main():
    async with stdio_server() as (read_stream, write_stream):
        await app.run(read_stream, write_stream, app.create_initialization_options())

if __name__ == "__main__":
    asyncio.run(main())

What Are MCP Resources and How Do They Differ from Tools?

Resources and tools serve different purposes in the MCP protocol:

Tools are actions the model can invoke. They execute operations and return results. Think of them as API endpoints the model can call.

Resources are data the model can read. They provide context without performing actions. Think of them as files or documents the model can access.

AspectToolsResources
PurposeExecute actionsProvide context
InvocationModel decides to callClient loads into context
Side effectsMay modify stateRead-only
Use caseQuery databases, run commandsDocumentation, configuration, state
Token costOnly when invokedLoaded upfront

Resources are particularly powerful for providing the model with organizational context — runbooks, architecture docs, coding standards — without consuming tool invocation budgets.

Production Deployment Patterns

Pattern 1: Sidecar Deployment

Run the MCP server as a sidecar container alongside the AI application. Shared network namespace eliminates latency.

Pattern 2: Gateway Pattern

A single MCP gateway server that routes tool invocations to appropriate backend services based on tool name prefixes. This simplifies client configuration — one connection instead of dozens.

Pattern 3: Tool Composition

MCP servers that call other MCP servers internally, composing higher-level tools from primitive operations. The outer server presents a simplified interface while orchestrating complex multi-step operations behind the scenes.

Performance Benchmarks for MCP Servers

Our production MCP servers handle 2.3 million invocations per month. Key performance metrics:

MetricP50P95P99
Tool discovery latency2ms8ms15ms
Tool invocation (simple)12ms34ms89ms
Tool invocation (DB query)45ms180ms420ms
Tool invocation (external API)120ms890ms2,400ms
Resource read (local file)4ms12ms28ms
Resource read (remote)35ms150ms380ms

Cost of MCP Infrastructure

ComponentMonthly Cost
Compute (3 server instances, 2 vCPU/4GB each)$180
Logging and monitoring$45
Secret management$12
Total$237
Cost per 1M invocations$103

MCP Server Latency Distribution by Tool Type

Security Best Practices for MCP Servers

  1. Input validation: Every tool input must be validated against its schema before execution. Never trust model-generated inputs blindly.

  2. Least privilege: MCP servers should have minimal permissions. A database tool server gets read-only access. A deployment tool gets deploy-only access.

  3. Rate limiting: Implement per-session and per-tool rate limits. A model stuck in a loop can generate thousands of tool calls per minute.

  4. Audit logging: Log every tool invocation with full arguments and results. This is essential for debugging and compliance.

  5. Sandboxing: Run tool execution in sandboxed environments. File system tools should be chrooted. Command execution should be containerized.

Key Takeaways

  • MCP standardizes AI tool integration, eliminating per-platform custom implementations and enabling tool portability across models
  • Three transport layers (stdio, SSE, streamable HTTP) serve different deployment needs from local tools to shared remote services
  • Tools execute actions; Resources provide context — understanding this distinction is critical for efficient server design
  • Production MCP servers handle millions of invocations with P50 latency under 50ms for most tool types
  • Infrastructure cost is minimal ($237/month for 2.3M invocations) making MCP economically viable at any scale
  • Security requires input validation, least privilege, rate limiting, audit logging, and sandboxed execution
  • The gateway pattern simplifies client configuration when managing many tool servers across an organization

Comments

    No comments yet. Be the first to share your thoughts.