You've built your first MCP server. It works on your laptop. The calculator adds numbers, the weather tool fetches data, everything seems perfect.
Then you deploy it to production and...
Users report "it just stopped working."
Error messages are cryptic: "Internal server error."
You have no idea what failed, when, or why.
Debugging takes hours because you can't see what's happening inside the server.
Welcome to the gap between prototype and production.
A production-ready MCP server isn't just about making tools work—it's about making them work reliably, observably, and maintainably at scale. This means proper async handling, comprehensive logging, structured error handling, and versioning strategies.
In this guide, you'll transform your prototype into a production-grade system that: handles thousands of concurrent requests, logs every important event, recovers gracefully from errors, and makes debugging a breeze! 🚀
🎯 The Production Gap: Why Prototypes Fail in Real World
Before diving into solutions, let's understand what breaks when you go from development to production.
The Five Production Killers
Your tool makes a 5-second API call.
During those 5 seconds, the entire server is frozen.
All other requests wait in line.
Result: Server becomes unusably slow under load.
An error occurs deep in your code.
No logs, no alerts, no visibility.
Users see "something went wrong" but you can't debug it.
Result: Hours wasted trying to reproduce issues.
Exception thrown:
KeyError: 'data'
No context about what operation failed or why.
Stack traces don't help without request context.
Result: Impossible to diagnose production issues.
You update a tool's parameters.
Old clients still call it with the old format.
Everything breaks without warning.
Result: Angry users and emergency rollbacks.
Your server opens database connections but never closes them.
Memory slowly leaks with each request.
After a few hours, server runs out of resources and crashes.
Result: Unpredictable outages and restarts.
Let's solve each of these systematically!
⚡ Async Tools: Handling Concurrent Requests
The first step to production readiness is making your tools asynchronous.
Why Async Matters
Synchronous (blocking) code:
def slow_tool():
time.sleep(5) # Blocks for 5 seconds
return "done"
# Request 1 arrives → waits 5 seconds
# Request 2 arrives → waits for Request 1, then waits 5 seconds
# Request 3 arrives → waits for 1 and 2, then waits 5 seconds
# Total time for 3 requests: 15 seconds!
Asynchronous (non-blocking) code:
async def fast_tool():
await asyncio.sleep(5) # Yields control
return "done"
# Request 1 arrives → starts waiting (non-blocking)
# Request 2 arrives → starts waiting (non-blocking)
# Request 3 arrives → starts waiting (non-blocking)
# All finish after ~5 seconds
# Total time for 3 requests: 5 seconds!
Async lets the server handle other requests while waiting for I/O operations.
This is critical because most MCP tools do I/O: API calls, database queries, file reads.
Without async, one slow request blocks everything!
Converting Sync to Async
Let's convert a synchronous tool to async step by step.
Before: Synchronous (BAD for production)
import requests
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Weather Server")
@mcp.tool()
def get_weather(city: str) -> dict:
"""Get weather for a city - BLOCKING VERSION"""
# This blocks the entire server!
response = requests.get(
f"https://api.weather.com/v1/current",
params={"city": city}
)
return response.json()
After: Asynchronous (GOOD for production)
import httpx # Async HTTP library
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Weather Server")
@mcp.tool()
async def get_weather(city: str) -> dict:
"""Get weather for a city - ASYNC VERSION"""
# This doesn't block other requests!
async with httpx.AsyncClient() as client:
response = await client.get(
f"https://api.weather.com/v1/current",
params={"city": city}
)
return response.json()
Common Async Patterns
Pattern 1: Async HTTP Calls
# ❌ Bad: Synchronous (blocks)
import requests
data = requests.get(url).json()
# ✅ Good: Asynchronous (non-blocking)
import httpx
async with httpx.AsyncClient() as client:
response = await client.get(url)
data = response.json()
Pattern 2: Async Database Queries
# ❌ Bad: Synchronous database call
import psycopg2
conn = psycopg2.connect(...)
cursor = conn.cursor()
cursor.execute("SELECT * FROM users")
results = cursor.fetchall()
# ✅ Good: Asynchronous database call
import asyncpg
conn = await asyncpg.connect(...)
results = await conn.fetch("SELECT * FROM users")
Pattern 3: Async File Operations
# ❌ Bad: Synchronous file read
with open("data.txt") as f:
content = f.read()
# ✅ Good: Asynchronous file read
import aiofiles
async with aiofiles.open("data.txt") as f:
content = await f.read()
Pattern 4: Parallel Async Operations
import asyncio
async def fetch_multiple_sources(cities: list) -> list:
"""Fetch weather for multiple cities in parallel"""
# Create tasks for all cities
tasks = [get_weather(city) for city in cities]
# Run all tasks concurrently
results = await asyncio.gather(*tasks)
return results
# Usage:
# Serial: 3 cities × 2 seconds each = 6 seconds total
# Parallel: max(2, 2, 2) = 2 seconds total!
• Add
async keyword to tool function
• Use
await for I/O operations
• Replace sync libraries with async versions (httpx, asyncpg, aiofiles)
• Use
asyncio.gather() for parallel operations
• Never use
time.sleep() — use await asyncio.sleep()
📊 Logging: Making Your Server Observable
Logging is your window into production. Without it, you're flying blind.
The Critical Logging Rule for STDIO Servers
For STDIO-based MCP servers (most common):
NEVER print to stdout!
print("Debug message") ← This will BREAK your server!
Why? STDIO servers communicate via stdout.
Your print statements corrupt the JSON-RPC messages.
Result: Client can't parse responses, everything fails.
Always log to stderr or files instead!
Setting Up Production Logging
Create production_logger.py:
"""
Production-Grade Logging Configuration
"""
import logging
import sys
import json
from datetime import datetime
from pathlib import Path
def setup_logger(name: str, log_level: str = "INFO") -> logging.Logger:
"""
Set up a production-ready logger
Features:
- Logs to stderr (safe for STDIO servers)
- Structured JSON format
- File logging with rotation
- Different levels for different environments
"""
logger = logging.getLogger(name)
logger.setLevel(getattr(logging, log_level.upper()))
# Remove existing handlers to avoid duplicates
logger.handlers.clear()
# ============================================
# STDERR Handler (Console)
# ============================================
stderr_handler = logging.StreamHandler(sys.stderr)
stderr_handler.setLevel(logging.INFO)
# Structured format with context
stderr_format = logging.Formatter(
fmt='%(asctime)s | %(levelname)-8s | %(name)s | %(message)s',
datefmt='%Y-%m-%d %H:%M:%S'
)
stderr_handler.setFormatter(stderr_format)
logger.addHandler(stderr_handler)
# ============================================
# File Handler (Persistent Logs)
# ============================================
log_dir = Path("logs")
log_dir.mkdir(exist_ok=True)
from logging.handlers import RotatingFileHandler
file_handler = RotatingFileHandler(
filename=log_dir / f"{name}.log",
maxBytes=10 * 1024 * 1024, # 10 MB
backupCount=5,
encoding='utf-8'
)
file_handler.setLevel(logging.DEBUG)
# JSON format for easy parsing
class JSONFormatter(logging.Formatter):
def format(self, record):
log_data = {
"timestamp": datetime.utcnow().isoformat(),
"level": record.levelname,
"logger": record.name,
"message": record.getMessage(),
"module": record.module,
"function": record.funcName,
"line": record.lineno
}
# Add exception info if present
if record.exc_info:
log_data["exception"] = self.formatException(record.exc_info)
# Add extra fields
if hasattr(record, 'extra'):
log_data.update(record.extra)
return json.dumps(log_data)
file_handler.setFormatter(JSONFormatter())
logger.addHandler(file_handler)
return logger
Using the Logger in Your Server
"""
Production MCP Server with Proper Logging
"""
from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger
import httpx
# Set up logger
logger = setup_logger("weather_server", log_level="INFO")
mcp = FastMCP("Weather Server")
@mcp.tool()
async def get_weather(city: str, ctx: Context) -> dict:
"""Get weather with comprehensive logging"""
# Log tool invocation
logger.info(f"Weather requested", extra={
"city": city,
"tool": "get_weather"
})
# Send progress to MCP client
await ctx.info(f"Fetching weather for {city}...")
try:
async with httpx.AsyncClient(timeout=10.0) as client:
logger.debug(f"Making API request", extra={
"city": city,
"endpoint": "api.weather.com"
})
response = await client.get(
"https://api.weather.com/v1/current",
params={"city": city}
)
response.raise_for_status()
data = response.json()
logger.info(f"Weather fetched successfully", extra={
"city": city,
"temperature": data.get("temp")
})
await ctx.info(f"Weather data retrieved successfully")
return {
"success": True,
"city": city,
"data": data
}
except httpx.TimeoutException:
logger.error(f"API timeout", extra={
"city": city,
"timeout": 10.0
})
await ctx.error(f"Request timed out after 10 seconds")
return {
"success": False,
"error": "API request timed out",
"city": city
}
except httpx.HTTPStatusError as e:
logger.error(f"HTTP error", extra={
"city": city,
"status_code": e.response.status_code,
"response": e.response.text
})
await ctx.error(f"API returned error: {e.response.status_code}")
return {
"success": False,
"error": f"API error: {e.response.status_code}",
"city": city
}
except Exception as e:
logger.exception(f"Unexpected error", extra={
"city": city,
"error_type": type(e).__name__
})
await ctx.error(f"Unexpected error occurred")
return {
"success": False,
"error": "Internal server error",
"city": city
}
MCP Client-Side Logging (Context Object)
MCP provides a special Context object that sends logs to the client.
Server-side logging (logger):
• Stored on server (files, databases)
• For debugging, auditing, monitoring
• Not visible to end users
Client-side logging (ctx):
• Sent to MCP client (Claude Desktop, etc.)
• User sees progress updates
• For transparency and user experience
Use BOTH!
@mcp.tool()
async def process_data(file_path: str, ctx: Context) -> dict:
"""Process data with dual logging"""
# Server-side: Technical details
logger.info("Data processing started", extra={
"file": file_path,
"user_id": ctx.request_id # Track request
})
# Client-side: User-friendly progress
await ctx.info("Starting data processing...")
# Log debug details (server only)
logger.debug(f"Reading file: {file_path}")
# Update user on progress
await ctx.info("Reading file...")
try:
# Process file...
await ctx.info("Processing complete!")
logger.info("Processing successful")
except ValueError as e:
# Server: Full details
logger.warning("Invalid data format", extra={
"file": file_path,
"error": str(e)
})
# Client: User-friendly message
await ctx.warning("File contains invalid data")
except Exception as e:
# Server: Full stack trace
logger.exception("Processing failed")
# Client: Safe error message
await ctx.error("Processing failed unexpectedly")
Log Levels: When to Use Each
| Level | When to Use | Example |
|---|---|---|
| DEBUG | Detailed info for diagnosing issues | Variable values, function calls |
| INFO | Normal operations, progress updates | Tool called, request completed |
| WARNING | Something unexpected but not fatal | Missing optional parameter, slow API |
| ERROR | Operation failed but server continues | API error, invalid input |
| CRITICAL | Severe error, server might crash | Database connection lost, out of memory |
🛡️ Error Handling: Graceful Failure
Production systems must handle errors gracefully—no cryptic messages, no crashes.
The Three-Layer Error Model
Structured Error Response Format
Always return errors in a consistent, parseable format:
class ErrorResponse:
"""Standardized error response structure"""
def __init__(
self,
error_code: str,
message: str,
details: dict = None,
retryable: bool = False
):
self.error_code = error_code
self.message = message
self.details = details or {}
self.retryable = retryable
def to_dict(self) -> dict:
return {
"success": False,
"error": {
"code": self.error_code,
"message": self.message,
"details": self.details,
"retryable": self.retryable
}
}
# Usage example:
return ErrorResponse(
error_code="API_TIMEOUT",
message="Weather API did not respond in time",
details={"timeout_seconds": 10, "city": city},
retryable=True
).to_dict()
Comprehensive Error Handling Pattern
import asyncio
from typing import Optional
import httpx
from mcp.server.fastmcp import FastMCP, Context
mcp = FastMCP("Robust Weather Server")
class WeatherAPIError(Exception):
"""Custom exception for weather API errors"""
pass
async def retry_with_backoff(
operation,
max_attempts: int = 3,
base_delay: float = 1.0
):
"""Retry failed operations with exponential backoff"""
for attempt in range(max_attempts):
try:
return await operation()
except Exception as e:
if attempt == max_attempts - 1:
# Last attempt failed, re-raise
raise
# Calculate delay with exponential backoff
delay = base_delay * (2 ** attempt)
logger.warning(
f"Attempt {attempt + 1} failed, retrying in {delay}s",
extra={"error": str(e)}
)
await asyncio.sleep(delay)
@mcp.tool()
async def get_weather_robust(
city: str,
ctx: Context,
units: str = "metric"
) -> dict:
"""
Get weather with comprehensive error handling
Returns structured responses for all scenarios:
- Success with data
- Retryable errors (network issues)
- Non-retryable errors (invalid input)
- Partial failures (degraded service)
"""
# ============================================
# STEP 1: Input Validation
# ============================================
if not city or not city.strip():
logger.warning("Empty city name provided")
await ctx.warning("City name cannot be empty")
return {
"success": False,
"error": {
"code": "INVALID_INPUT",
"message": "City name is required",
"details": {"parameter": "city"},
"retryable": False
}
}
if units not in ["metric", "imperial"]:
logger.warning(f"Invalid units: {units}")
await ctx.warning(f"Invalid units '{units}', using 'metric'")
units = "metric" # Fallback to default
# ============================================
# STEP 2: Execute with Retry Logic
# ============================================
logger.info(f"Fetching weather", extra={
"city": city,
"units": units
})
await ctx.info(f"Fetching weather for {city}...")
try:
async def fetch_weather():
async with httpx.AsyncClient(timeout=10.0) as client:
response = await client.get(
"https://api.weather.com/v1/current",
params={"city": city, "units": units}
)
response.raise_for_status()
return response.json()
# Retry up to 3 times with backoff
data = await retry_with_backoff(fetch_weather, max_attempts=3)
logger.info("Weather fetched successfully", extra={
"city": city,
"temp": data.get("temperature")
})
await ctx.info("Weather data retrieved")
return {
"success": True,
"city": city,
"units": units,
"data": data,
"cached": False
}
# ============================================
# STEP 3: Handle Specific Errors
# ============================================
except httpx.TimeoutException as e:
logger.error("API timeout", extra={
"city": city,
"timeout": 10.0
})
await ctx.error("Weather service timed out")
return {
"success": False,
"error": {
"code": "API_TIMEOUT",
"message": "Weather service did not respond in time",
"details": {
"city": city,
"timeout_seconds": 10.0
},
"retryable": True
}
}
except httpx.HTTPStatusError as e:
status_code = e.response.status_code
if status_code == 404:
logger.warning(f"City not found", extra={
"city": city,
"status": 404
})
await ctx.warning(f"City '{city}' not found")
return {
"success": False,
"error": {
"code": "CITY_NOT_FOUND",
"message": f"Weather data not available for '{city}'",
"details": {
"city": city,
"suggestion": "Check city name spelling"
},
"retryable": False
}
}
elif status_code == 429:
logger.error("Rate limit exceeded", extra={
"city": city,
"status": 429
})
await ctx.error("Too many requests, please try again later")
return {
"success": False,
"error": {
"code": "RATE_LIMIT",
"message": "Too many requests to weather service",
"details": {
"retry_after_seconds": 60
},
"retryable": True
}
}
else:
logger.error(f"HTTP error", extra={
"city": city,
"status": status_code,
"response": e.response.text[:200]
})
await ctx.error(f"Weather service error: {status_code}")
return {
"success": False,
"error": {
"code": "API_ERROR",
"message": f"Weather service returned error {status_code}",
"details": {"status_code": status_code},
"retryable": status_code >= 500
}
}
except httpx.NetworkError as e:
logger.error("Network error", extra={
"city": city,
"error": str(e)
})
await ctx.error("Network connection failed")
return {
"success": False,
"error": {
"code": "NETWORK_ERROR",
"message": "Could not connect to weather service",
"details": {"error": str(e)},
"retryable": True
}
}
except Exception as e:
# Catch-all for unexpected errors
logger.exception("Unexpected error", extra={
"city": city,
"error_type": type(e).__name__
})
await ctx.error("An unexpected error occurred")
return {
"success": False,
"error": {
"code": "INTERNAL_ERROR",
"message": "An unexpected error occurred",
"details": {
"error_type": type(e).__name__
},
"retryable": False
}
}
• Always return structured error objects (not just strings)
• Include error codes for programmatic handling
• Mark errors as retryable or not
• Provide actionable error messages
• Log full details server-side, safe messages client-side
• Use retry logic for transient failures
• Never expose internal implementation details in errors
🔖 Versioning: Managing Breaking Changes
As your server evolves, you'll need to update tools without breaking existing clients.
Versioning Strategies
Strategy 1: Semantic Versioning in Server Name
from mcp.server.fastmcp import FastMCP
# Version embedded in server name
mcp = FastMCP("Weather Server v2.1.0")
# Expose version as resource
@mcp.resource("version://info")
def get_version() -> dict:
return {
"version": "2.1.0",
"released": "2024-02-15",
"breaking_changes": [
"Removed deprecated 'temp_fahrenheit' field",
"Changed 'location' parameter to 'city'"
],
"deprecated": [
"'units' parameter will be required in v3.0.0"
]
}
Strategy 2: Versioned Tool Names
# Support multiple versions simultaneously
@mcp.tool()
async def get_weather_v1(location: str) -> dict:
"""
DEPRECATED: Use get_weather_v2 instead
Will be removed in version 3.0.0
"""
logger.warning("get_weather_v1 called - deprecated")
# Convert to new format internally
return await get_weather_v2(city=location)
@mcp.tool()
async def get_weather_v2(city: str, units: str = "metric") -> dict:
"""
Current version: Get weather data
Args:
city: City name (replaces 'location' from v1)
units: Temperature units (new in v2)
"""
# Implementation...
Strategy 3: Parameter Evolution (Backward Compatible)
from typing import Optional
@mcp.tool()
async def get_weather(
city: Optional[str] = None,
location: Optional[str] = None, # Deprecated, kept for compatibility
units: str = "metric"
) -> dict:
"""
Get weather data
Args:
city: City name (preferred)
location: DEPRECATED - use 'city' instead
units: Temperature units
"""
# Handle both old and new parameter names
if location and not city:
logger.warning("'location' parameter deprecated, use 'city'")
city = location
if not city:
return {
"success": False,
"error": {
"code": "MISSING_PARAMETER",
"message": "Either 'city' or 'location' is required"
}
}
# Rest of implementation...
Deprecation Workflow
• Give users 6-12 months notice before breaking changes
• Support old + new simultaneously during transition
• Log deprecation warnings prominently
• Provide migration guides and tools
• Use semantic versioning: MAJOR.MINOR.PATCH
• Increment MAJOR for breaking changes only
🎯 Complete Production-Ready Example
Let's put it all together! Here's a production-grade MCP server.
Create production_weather_server.py:
"""
Production-Ready Weather MCP Server
Features:
- Async tools for concurrency
- Comprehensive logging (server + client)
- Structured error handling
- Retry logic with backoff
- Versioning support
- Resource monitoring
"""
import asyncio
import sys
from datetime import datetime
from typing import Optional
import httpx
from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger
# ============================================
# CONFIGURATION
# ============================================
SERVER_VERSION = "2.1.0"
API_TIMEOUT = 10.0
MAX_RETRIES = 3
# ============================================
# LOGGING SETUP
# ============================================
logger = setup_logger("weather_server", log_level="INFO")
# ============================================
# SERVER INITIALIZATION
# ============================================
mcp = FastMCP(f"Weather Server v{SERVER_VERSION}")
logger.info("Server starting", extra={
"version": SERVER_VERSION,
"python_version": sys.version,
"timestamp": datetime.utcnow().isoformat()
})
# ============================================
# UTILITY FUNCTIONS
# ============================================
async def retry_operation(operation, max_attempts: int = MAX_RETRIES):
"""Retry operation with exponential backoff"""
for attempt in range(max_attempts):
try:
return await operation()
except Exception as e:
if attempt == max_attempts - 1:
raise
delay = (2 ** attempt) + (asyncio.get_event_loop().time() % 1)
logger.debug(f"Retry {attempt + 1}/{max_attempts}", extra={
"delay": delay,
"error": str(e)
})
await asyncio.sleep(delay)
def sanitize_city_name(city: str) -> str:
"""Clean and validate city name"""
return city.strip().title()
# ============================================
# RESOURCE: Server Information
# ============================================
@mcp.resource("server://info")
def server_info() -> dict:
"""Get server version and status information"""
return {
"version": SERVER_VERSION,
"name": "Weather Server",
"status": "operational",
"features": [
"async_tools",
"retry_logic",
"structured_errors",
"client_logging"
],
"api_timeout_seconds": API_TIMEOUT,
"max_retries": MAX_RETRIES
}
# ============================================
# TOOL: Get Weather (Current Version)
# ============================================
@mcp.tool()
async def get_weather(
city: str,
ctx: Context,
units: str = "metric"
) -> dict:
"""
Get current weather for a city
Args:
city: City name (e.g., "London", "New York")
units: Temperature units - "metric" (Celsius) or "imperial" (Fahrenheit)
Returns:
Weather data with temperature, conditions, humidity, etc.
"""
request_id = id(ctx)
logger.info("Tool invoked", extra={
"tool": "get_weather",
"city": city,
"units": units,
"request_id": request_id
})
# Validate inputs
if not city or not city.strip():
logger.warning("Invalid input: empty city", extra={
"request_id": request_id
})
await ctx.warning("City name cannot be empty")
return {
"success": False,
"error": {
"code": "INVALID_INPUT",
"message": "City name is required",
"retryable": False
}
}
city = sanitize_city_name(city)
if units not in ["metric", "imperial"]:
logger.warning(f"Invalid units: {units}", extra={
"request_id": request_id
})
await ctx.warning(f"Invalid units '{units}', using 'metric'")
units = "metric"
# Fetch weather data
await ctx.info(f"Fetching weather for {city}...")
try:
async def fetch():
async with httpx.AsyncClient(timeout=API_TIMEOUT) as client:
logger.debug("API request starting", extra={
"city": city,
"request_id": request_id
})
response = await client.get(
f"https://api.weather.com/v1/current",
params={"city": city, "units": units}
)
response.raise_for_status()
return response.json()
data = await retry_operation(fetch, max_attempts=MAX_RETRIES)
logger.info("Weather retrieved successfully", extra={
"city": city,
"temperature": data.get("temperature"),
"request_id": request_id
})
await ctx.info("Weather data retrieved")
return {
"success": True,
"city": city,
"units": units,
"data": data,
"server_version": SERVER_VERSION,
"timestamp": datetime.utcnow().isoformat()
}
except httpx.TimeoutException:
logger.error("API timeout", extra={
"city": city,
"timeout": API_TIMEOUT,
"request_id": request_id
})
await ctx.error(f"Request timed out after {API_TIMEOUT} seconds")
return {
"success": False,
"error": {
"code": "API_TIMEOUT",
"message": "Weather service did not respond in time",
"details": {"timeout_seconds": API_TIMEOUT},
"retryable": True
}
}
except httpx.HTTPStatusError as e:
status = e.response.status_code
logger.error("HTTP error", extra={
"city": city,
"status_code": status,
"request_id": request_id
})
if status == 404:
await ctx.warning(f"City '{city}' not found")
return {
"success": False,
"error": {
"code": "CITY_NOT_FOUND",
"message": f"No weather data for '{city}'",
"retryable": False
}
}
await ctx.error(f"Weather service error: {status}")
return {
"success": False,
"error": {
"code": "API_ERROR",
"message": f"HTTP {status}",
"retryable": status >= 500
}
}
except Exception as e:
logger.exception("Unexpected error", extra={
"city": city,
"request_id": request_id
})
await ctx.error("An unexpected error occurred")
return {
"success": False,
"error": {
"code": "INTERNAL_ERROR",
"message": "Unexpected error",
"retryable": False
}
}
# ============================================
# RUN SERVER
# ============================================
if __name__ == "__main__":
logger.info("Server ready to accept connections")
print(f" Weather Server v{SERVER_VERSION} starting...", file=sys.stderr)
print(f"Logging to: logs/weather_server.log", file=sys.stderr)
print(f" Server operational\n", file=sys.stderr)
try:
mcp.run()
except KeyboardInterrupt:
logger.info("Server shutting down (KeyboardInterrupt)")
print("\n Server shutdown", file=sys.stderr)
except Exception as e:
logger.critical("Server crash", extra={
"error": str(e)
}, exc_info=True)
print(f"\n Server crashed: {e}", file=sys.stderr)
sys.exit(1)
🧪 Exercise: Add Logging and Error Handling
Time to apply what you've learned! Take your existing MCP server and make it production-ready.
Exercise Instructions
Take the calculator server from previous lessons and add:
1. Production logging (server + client)
2. Structured error handling
3. Async tool implementation
4. Version information
Starter Code
# Before (basic prototype)
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Calculator")
@mcp.tool()
def add(a: float, b: float) -> float:
return a + b
@mcp.tool()
def divide(a: float, b: float) -> float:
return a / b # What if b is zero?
if __name__ == "__main__":
mcp.run()
Solution: Production-Ready Calculator
# After (production-ready)
"""
Production-Ready Calculator MCP Server
"""
from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger
from typing import Union
import sys
# Configuration
VERSION = "1.0.0"
# Setup logging
logger = setup_logger("calculator_server", log_level="INFO")
# Create server
mcp = FastMCP(f"Calculator v{VERSION}")
logger.info("Calculator server initializing", extra={
"version": VERSION
})
# ============================================
# RESOURCE: Server Info
# ============================================
@mcp.resource("calculator://version")
def get_version() -> dict:
"""Get calculator version information"""
return {
"version": VERSION,
"operations": ["add", "subtract", "multiply", "divide"],
"status": "operational"
}
# ============================================
# TOOL: Add Numbers
# ============================================
@mcp.tool()
async def add(a: float, b: float, ctx: Context) -> dict:
"""
Add two numbers with perfect precision
Args:
a: First number
b: Second number
"""
logger.info("Add operation", extra={
"a": a,
"b": b
})
await ctx.info(f"Adding {a} + {b}...")
try:
result = a + b
logger.info("Addition successful", extra={
"result": result
})
await ctx.info(f"Result: {result}")
return {
"success": True,
"operation": "addition",
"a": a,
"b": b,
"result": result,
"formula": f"{a} + {b} = {result}"
}
except Exception as e:
logger.exception("Addition failed", extra={
"a": a,
"b": b
})
await ctx.error("Addition failed unexpectedly")
return {
"success": False,
"error": {
"code": "CALCULATION_ERROR",
"message": "Could not perform addition",
"retryable": False
}
}
# ============================================
# TOOL: Divide Numbers (with error handling)
# ============================================
@mcp.tool()
async def divide(a: float, b: float, ctx: Context) -> dict:
"""
Divide two numbers safely
Args:
a: Numerator
b: Denominator (cannot be zero)
"""
logger.info("Divide operation", extra={
"a": a,
"b": b
})
await ctx.info(f"Dividing {a} ÷ {b}...")
# Handle division by zero
if b == 0:
logger.warning("Division by zero attempted", extra={
"a": a,
"b": b
})
await ctx.warning("Cannot divide by zero")
return {
"success": False,
"error": {
"code": "DIVISION_BY_ZERO",
"message": "Cannot divide by zero",
"details": {
"numerator": a,
"denominator": b
},
"retryable": False
}
}
try:
result = a / b
logger.info("Division successful", extra={
"result": result
})
await ctx.info(f"Result: {result}")
return {
"success": True,
"operation": "division",
"a": a,
"b": b,
"result": result,
"formula": f"{a} ÷ {b} = {result}"
}
except Exception as e:
logger.exception("Division failed", extra={
"a": a,
"b": b
})
await ctx.error("Division failed unexpectedly")
return {
"success": False,
"error": {
"code": "CALCULATION_ERROR",
"message": "Could not perform division",
"retryable": False
}
}
# ============================================
# MAIN
# ============================================
if __name__ == "__main__":
logger.info("Server ready")
print(f" Calculator Server v{VERSION} starting...", file=sys.stderr)
print(f" Logging to: logs/calculator_server.log", file=sys.stderr)
print(f" Server operational\n", file=sys.stderr)
try:
mcp.run()
except KeyboardInterrupt:
logger.info("Server shutdown")
print("\n Goodbye!", file=sys.stderr)
Testing Your Production Server
# Run server
python production_calculator_server.py
# In another terminal, test with MCP Inspector
mcp dev production_calculator_server.py
# Try these test cases:
# 1. Normal: add(5, 3) → Should succeed
# 2. Error: divide(10, 0) → Should return structured error
# 3. Check logs/calculator_server.log for detailed logging
• Async functions (async def)
• Server-side logging (logger.info, logger.error)
• Client-side logging (ctx.info, ctx.error)
• Structured error responses
• Version resource
• Graceful error handling (division by zero)
Production Checklist
Before deploying your MCP server to production, use this checklist:
□ All I/O operations are async
□ No blocking calls (time.sleep, requests.get, etc.)
□ Proper connection pooling
□ Timeouts on all external calls
□ Retry logic for transient failures
□ Never print to stdout (STDIO servers)
□ Structured logging format (JSON)
□ Log levels properly used
□ Client-side progress updates (ctx)
□ Sensitive data sanitized from logs
□ Log rotation configured
□ All errors return structured responses
□ Errors include error codes
□ Retryable flag set correctly
□ User-friendly error messages
□ No internal details leaked
□ Proper exception logging
□ Server version exposed
□ Breaking changes documented
□ Deprecation warnings in place
□ Backward compatibility maintained
□ Migration guides provided
□ All inputs validated
□ SQL injection prevention
□ Path traversal protection
□ Rate limiting implemented
□ Authentication/authorization
🚀 Advanced Production Topics
Monitoring & Metrics
import time
from prometheus_client import Counter, Histogram
# Track tool usage
tool_calls = Counter('mcp_tool_calls_total', 'Total tool calls', ['tool_name', 'status'])
tool_duration = Histogram('mcp_tool_duration_seconds', 'Tool execution time', ['tool_name'])
@mcp.tool()
async def monitored_tool(ctx: Context) -> dict:
start = time.time()
tool_name = "monitored_tool"
try:
result = await do_work()
tool_calls.labels(tool_name=tool_name, status='success').inc()
return result
except Exception as e:
tool_calls.labels(tool_name=tool_name, status='error').inc()
raise
finally:
duration = time.time() - start
tool_duration.labels(tool_name=tool_name).observe(duration)
Health Checks
@mcp.resource("health://status")
async def health_check() -> dict:
"""Server health status"""
checks = {
"database": await check_database(),
"external_api": await check_external_api(),
"disk_space": check_disk_space()
}
healthy = all(checks.values())
return {
"status": "healthy" if healthy else "degraded",
"checks": checks,
"timestamp": datetime.utcnow().isoformat()
}
Graceful Shutdown
import signal
import asyncio
shutdown_event = asyncio.Event()
def handle_shutdown(signum, frame):
logger.info("Shutdown signal received")
shutdown_event.set()
signal.signal(signal.SIGTERM, handle_shutdown)
signal.signal(signal.SIGINT, handle_shutdown)
async def cleanup():
"""Cleanup resources before shutdown"""
logger.info("Performing cleanup...")
await close_database_connections()
await flush_logs()
logger.info("Cleanup complete")
if __name__ == "__main__":
try:
mcp.run()
finally:
asyncio.run(cleanup())
📚 Summary & Next Steps
• Async tools prevent blocking and enable concurrency
• Proper logging (stderr, not stdout) enables debugging
• Structured errors help clients handle failures gracefully
• Versioning prevents breaking changes from disrupting users
• Production readiness is about reliability, not just functionality
Week 1: Convert all your tools to async
Week 2: Add comprehensive logging to every tool
Week 3: Implement structured error handling
Week 4: Add versioning and monitoring
Week 5: Load test and optimize
Essential Resources
• httpx - Async HTTP client
• asyncpg - Async PostgreSQL
• aiofiles - Async file operations
• structlog - Structured logging
Documentation:
• Python AsyncIO: docs.python.org/3/library/asyncio
• MCP Specification: modelcontextprotocol.io
• Logging Best Practices: 12factor.net/logs
✨ Final Thoughts
The difference between a prototype and production isn't complexity—it's reliability.
A prototype works when everything goes right.
A production system works especially when things go wrong.
You've learned:
- How to handle thousands of concurrent requests efficiently
- How to debug issues in production without guessing
- How to fail gracefully and communicate errors clearly
- How to evolve your server without breaking clients
These aren't "nice to haves"—they're requirements for any system users depend on.
Take your existing MCP server (calculator, weather, or custom).
Spend this week making it production-ready using this guide.
Then deploy it somewhere (local, cloud, anywhere) and use it for real work.
You'll discover edge cases you never imagined.
You'll appreciate good logging when debugging at 2 AM.
You'll thank yourself for structured errors when users report issues.
Production readiness is a journey, not a destination.
Start with async. Add logging. Handle errors. Version carefully.
Monitor everything. Improve iteratively.
Comments
Post a Comment