Skip to main content

Production-Ready MCP Server

Calculating read time…

You've built your first MCP server. It works on your laptop. The calculator adds numbers, the weather tool fetches data, everything seems perfect.

Then you deploy it to production and...

Users report "it just stopped working."
Error messages are cryptic: "Internal server error."
You have no idea what failed, when, or why.
Debugging takes hours because you can't see what's happening inside the server.

Welcome to the gap between prototype and production.

A production-ready MCP server isn't just about making tools work—it's about making them work reliably, observably, and maintainably at scale. This means proper async handling, comprehensive logging, structured error handling, and versioning strategies.

In this guide, you'll transform your prototype into a production-grade system that: handles thousands of concurrent requests, logs every important event, recovers gracefully from errors, and makes debugging a breeze! 🚀

🎯 The Production Gap: Why Prototypes Fail in Real World

Before diving into solutions, let's understand what breaks when you go from development to production.

The Five Production Killers

❌ Killer #1: Blocking Operations

Your tool makes a 5-second API call.
During those 5 seconds, the entire server is frozen.
All other requests wait in line.
Result: Server becomes unusably slow under load.
❌ Killer #2: Silent Failures

An error occurs deep in your code.
No logs, no alerts, no visibility.
Users see "something went wrong" but you can't debug it.
Result: Hours wasted trying to reproduce issues.
❌ Killer #3: Cryptic Errors

Exception thrown: KeyError: 'data'
No context about what operation failed or why.
Stack traces don't help without request context.
Result: Impossible to diagnose production issues.
❌ Killer #4: Breaking Changes

You update a tool's parameters.
Old clients still call it with the old format.
Everything breaks without warning.
Result: Angry users and emergency rollbacks.
❌ Killer #5: Uncontrolled Resource Usage

Your server opens database connections but never closes them.
Memory slowly leaks with each request.
After a few hours, server runs out of resources and crashes.
Result: Unpredictable outages and restarts.

Let's solve each of these systematically!

⚡ Async Tools: Handling Concurrent Requests

The first step to production readiness is making your tools asynchronous.

Why Async Matters

Synchronous (blocking) code:

def slow_tool():
    time.sleep(5)  # Blocks for 5 seconds
    return "done"

# Request 1 arrives → waits 5 seconds
# Request 2 arrives → waits for Request 1, then waits 5 seconds
# Request 3 arrives → waits for 1 and 2, then waits 5 seconds
# Total time for 3 requests: 15 seconds!

Asynchronous (non-blocking) code:

async def fast_tool():
    await asyncio.sleep(5)  # Yields control
    return "done"

# Request 1 arrives → starts waiting (non-blocking)
# Request 2 arrives → starts waiting (non-blocking)
# Request 3 arrives → starts waiting (non-blocking)
# All finish after ~5 seconds
# Total time for 3 requests: 5 seconds!
🎓 The Key Insight:

Async lets the server handle other requests while waiting for I/O operations.
This is critical because most MCP tools do I/O: API calls, database queries, file reads.
Without async, one slow request blocks everything!

Converting Sync to Async

Let's convert a synchronous tool to async step by step.

Before: Synchronous (BAD for production)

import requests
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("Weather Server")

@mcp.tool()
def get_weather(city: str) -> dict:
    """Get weather for a city - BLOCKING VERSION"""
    # This blocks the entire server!
    response = requests.get(
        f"https://api.weather.com/v1/current",
        params={"city": city}
    )
    return response.json()

After: Asynchronous (GOOD for production)

import httpx  # Async HTTP library
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("Weather Server")

@mcp.tool()
async def get_weather(city: str) -> dict:
    """Get weather for a city - ASYNC VERSION"""
    # This doesn't block other requests!
    async with httpx.AsyncClient() as client:
        response = await client.get(
            f"https://api.weather.com/v1/current",
            params={"city": city}
        )
        return response.json()

Common Async Patterns

Pattern 1: Async HTTP Calls

# ❌ Bad: Synchronous (blocks)
import requests
data = requests.get(url).json()

# ✅ Good: Asynchronous (non-blocking)
import httpx
async with httpx.AsyncClient() as client:
    response = await client.get(url)
    data = response.json()

Pattern 2: Async Database Queries

# ❌ Bad: Synchronous database call
import psycopg2
conn = psycopg2.connect(...)
cursor = conn.cursor()
cursor.execute("SELECT * FROM users")
results = cursor.fetchall()

# ✅ Good: Asynchronous database call
import asyncpg
conn = await asyncpg.connect(...)
results = await conn.fetch("SELECT * FROM users")

Pattern 3: Async File Operations

# ❌ Bad: Synchronous file read
with open("data.txt") as f:
    content = f.read()

# ✅ Good: Asynchronous file read
import aiofiles
async with aiofiles.open("data.txt") as f:
    content = await f.read()

Pattern 4: Parallel Async Operations

import asyncio

async def fetch_multiple_sources(cities: list) -> list:
    """Fetch weather for multiple cities in parallel"""
    
    # Create tasks for all cities
    tasks = [get_weather(city) for city in cities]
    
    # Run all tasks concurrently
    results = await asyncio.gather(*tasks)
    
    return results

# Usage:
# Serial: 3 cities × 2 seconds each = 6 seconds total
# Parallel: max(2, 2, 2) = 2 seconds total!
✅ Async Tool Checklist:

• Add async keyword to tool function
• Use await for I/O operations
• Replace sync libraries with async versions (httpx, asyncpg, aiofiles)
• Use asyncio.gather() for parallel operations
• Never use time.sleep() — use await asyncio.sleep()

📊 Logging: Making Your Server Observable

Logging is your window into production. Without it, you're flying blind.

The Critical Logging Rule for STDIO Servers

⚠️ CRITICAL WARNING:

For STDIO-based MCP servers (most common):
NEVER print to stdout!

print("Debug message") ← This will BREAK your server!

Why? STDIO servers communicate via stdout.
Your print statements corrupt the JSON-RPC messages.
Result: Client can't parse responses, everything fails.

Always log to stderr or files instead!

Setting Up Production Logging

Create production_logger.py:

"""
Production-Grade Logging Configuration
"""

import logging
import sys
import json
from datetime import datetime
from pathlib import Path

def setup_logger(name: str, log_level: str = "INFO") -> logging.Logger:
    """
    Set up a production-ready logger
    
    Features:
    - Logs to stderr (safe for STDIO servers)
    - Structured JSON format
    - File logging with rotation
    - Different levels for different environments
    """
    
    logger = logging.getLogger(name)
    logger.setLevel(getattr(logging, log_level.upper()))
    
    # Remove existing handlers to avoid duplicates
    logger.handlers.clear()
    
    # ============================================
    # STDERR Handler (Console)
    # ============================================
    
    stderr_handler = logging.StreamHandler(sys.stderr)
    stderr_handler.setLevel(logging.INFO)
    
    # Structured format with context
    stderr_format = logging.Formatter(
        fmt='%(asctime)s | %(levelname)-8s | %(name)s | %(message)s',
        datefmt='%Y-%m-%d %H:%M:%S'
    )
    stderr_handler.setFormatter(stderr_format)
    logger.addHandler(stderr_handler)
    
    # ============================================
    # File Handler (Persistent Logs)
    # ============================================
    
    log_dir = Path("logs")
    log_dir.mkdir(exist_ok=True)
    
    from logging.handlers import RotatingFileHandler
    
    file_handler = RotatingFileHandler(
        filename=log_dir / f"{name}.log",
        maxBytes=10 * 1024 * 1024,  # 10 MB
        backupCount=5,
        encoding='utf-8'
    )
    file_handler.setLevel(logging.DEBUG)
    
    # JSON format for easy parsing
    class JSONFormatter(logging.Formatter):
        def format(self, record):
            log_data = {
                "timestamp": datetime.utcnow().isoformat(),
                "level": record.levelname,
                "logger": record.name,
                "message": record.getMessage(),
                "module": record.module,
                "function": record.funcName,
                "line": record.lineno
            }
            
            # Add exception info if present
            if record.exc_info:
                log_data["exception"] = self.formatException(record.exc_info)
            
            # Add extra fields
            if hasattr(record, 'extra'):
                log_data.update(record.extra)
            
            return json.dumps(log_data)
    
    file_handler.setFormatter(JSONFormatter())
    logger.addHandler(file_handler)
    
    return logger

Using the Logger in Your Server

"""
Production MCP Server with Proper Logging
"""

from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger
import httpx

# Set up logger
logger = setup_logger("weather_server", log_level="INFO")

mcp = FastMCP("Weather Server")

@mcp.tool()
async def get_weather(city: str, ctx: Context) -> dict:
    """Get weather with comprehensive logging"""
    
    # Log tool invocation
    logger.info(f"Weather requested", extra={
        "city": city,
        "tool": "get_weather"
    })
    
    # Send progress to MCP client
    await ctx.info(f"Fetching weather for {city}...")
    
    try:
        async with httpx.AsyncClient(timeout=10.0) as client:
            logger.debug(f"Making API request", extra={
                "city": city,
                "endpoint": "api.weather.com"
            })
            
            response = await client.get(
                "https://api.weather.com/v1/current",
                params={"city": city}
            )
            
            response.raise_for_status()
            data = response.json()
            
            logger.info(f"Weather fetched successfully", extra={
                "city": city,
                "temperature": data.get("temp")
            })
            
            await ctx.info(f"Weather data retrieved successfully")
            
            return {
                "success": True,
                "city": city,
                "data": data
            }
            
    except httpx.TimeoutException:
        logger.error(f"API timeout", extra={
            "city": city,
            "timeout": 10.0
        })
        await ctx.error(f"Request timed out after 10 seconds")
        
        return {
            "success": False,
            "error": "API request timed out",
            "city": city
        }
        
    except httpx.HTTPStatusError as e:
        logger.error(f"HTTP error", extra={
            "city": city,
            "status_code": e.response.status_code,
            "response": e.response.text
        })
        await ctx.error(f"API returned error: {e.response.status_code}")
        
        return {
            "success": False,
            "error": f"API error: {e.response.status_code}",
            "city": city
        }
        
    except Exception as e:
        logger.exception(f"Unexpected error", extra={
            "city": city,
            "error_type": type(e).__name__
        })
        await ctx.error(f"Unexpected error occurred")
        
        return {
            "success": False,
            "error": "Internal server error",
            "city": city
        }

MCP Client-Side Logging (Context Object)

MCP provides a special Context object that sends logs to the client.

🎓 Two Types of Logging:

Server-side logging (logger):
• Stored on server (files, databases)
• For debugging, auditing, monitoring
• Not visible to end users

Client-side logging (ctx):
• Sent to MCP client (Claude Desktop, etc.)
• User sees progress updates
• For transparency and user experience

Use BOTH!
@mcp.tool()
async def process_data(file_path: str, ctx: Context) -> dict:
    """Process data with dual logging"""
    
    # Server-side: Technical details
    logger.info("Data processing started", extra={
        "file": file_path,
        "user_id": ctx.request_id  # Track request
    })
    
    # Client-side: User-friendly progress
    await ctx.info("Starting data processing...")
    
    # Log debug details (server only)
    logger.debug(f"Reading file: {file_path}")
    
    # Update user on progress
    await ctx.info("Reading file...")
    
    try:
        # Process file...
        await ctx.info("Processing complete!")
        logger.info("Processing successful")
        
    except ValueError as e:
        # Server: Full details
        logger.warning("Invalid data format", extra={
            "file": file_path,
            "error": str(e)
        })
        # Client: User-friendly message
        await ctx.warning("File contains invalid data")
        
    except Exception as e:
        # Server: Full stack trace
        logger.exception("Processing failed")
        # Client: Safe error message
        await ctx.error("Processing failed unexpectedly")

Log Levels: When to Use Each

Level When to Use Example
DEBUG Detailed info for diagnosing issues Variable values, function calls
INFO Normal operations, progress updates Tool called, request completed
WARNING Something unexpected but not fatal Missing optional parameter, slow API
ERROR Operation failed but server continues API error, invalid input
CRITICAL Severe error, server might crash Database connection lost, out of memory

🛡️ Error Handling: Graceful Failure

Production systems must handle errors gracefully—no cryptic messages, no crashes.

The Three-Layer Error Model

Layer 1: TRANSPORT ERRORS ├─ Network failures ├─ Connection timeouts └─ Authentication errors Layer 2: PROTOCOL ERRORS (JSON-RPC) ├─ Malformed JSON ├─ Invalid method calls └─ Missing parameters Layer 3: APPLICATION ERRORS (Your Code) ├─ Business logic failures ├─ External API errors └─ Resource constraints

Structured Error Response Format

Always return errors in a consistent, parseable format:

class ErrorResponse:
    """Standardized error response structure"""
    
    def __init__(
        self,
        error_code: str,
        message: str,
        details: dict = None,
        retryable: bool = False
    ):
        self.error_code = error_code
        self.message = message
        self.details = details or {}
        self.retryable = retryable
    
    def to_dict(self) -> dict:
        return {
            "success": False,
            "error": {
                "code": self.error_code,
                "message": self.message,
                "details": self.details,
                "retryable": self.retryable
            }
        }

# Usage example:
return ErrorResponse(
    error_code="API_TIMEOUT",
    message="Weather API did not respond in time",
    details={"timeout_seconds": 10, "city": city},
    retryable=True
).to_dict()

Comprehensive Error Handling Pattern

import asyncio
from typing import Optional
import httpx
from mcp.server.fastmcp import FastMCP, Context

mcp = FastMCP("Robust Weather Server")

class WeatherAPIError(Exception):
    """Custom exception for weather API errors"""
    pass

async def retry_with_backoff(
    operation,
    max_attempts: int = 3,
    base_delay: float = 1.0
):
    """Retry failed operations with exponential backoff"""
    
    for attempt in range(max_attempts):
        try:
            return await operation()
        except Exception as e:
            if attempt == max_attempts - 1:
                # Last attempt failed, re-raise
                raise
            
            # Calculate delay with exponential backoff
            delay = base_delay * (2 ** attempt)
            logger.warning(
                f"Attempt {attempt + 1} failed, retrying in {delay}s",
                extra={"error": str(e)}
            )
            await asyncio.sleep(delay)

@mcp.tool()
async def get_weather_robust(
    city: str,
    ctx: Context,
    units: str = "metric"
) -> dict:
    """
    Get weather with comprehensive error handling
    
    Returns structured responses for all scenarios:
    - Success with data
    - Retryable errors (network issues)
    - Non-retryable errors (invalid input)
    - Partial failures (degraded service)
    """
    
    # ============================================
    # STEP 1: Input Validation
    # ============================================
    
    if not city or not city.strip():
        logger.warning("Empty city name provided")
        await ctx.warning("City name cannot be empty")
        
        return {
            "success": False,
            "error": {
                "code": "INVALID_INPUT",
                "message": "City name is required",
                "details": {"parameter": "city"},
                "retryable": False
            }
        }
    
    if units not in ["metric", "imperial"]:
        logger.warning(f"Invalid units: {units}")
        await ctx.warning(f"Invalid units '{units}', using 'metric'")
        units = "metric"  # Fallback to default
    
    # ============================================
    # STEP 2: Execute with Retry Logic
    # ============================================
    
    logger.info(f"Fetching weather", extra={
        "city": city,
        "units": units
    })
    await ctx.info(f"Fetching weather for {city}...")
    
    try:
        async def fetch_weather():
            async with httpx.AsyncClient(timeout=10.0) as client:
                response = await client.get(
                    "https://api.weather.com/v1/current",
                    params={"city": city, "units": units}
                )
                response.raise_for_status()
                return response.json()
        
        # Retry up to 3 times with backoff
        data = await retry_with_backoff(fetch_weather, max_attempts=3)
        
        logger.info("Weather fetched successfully", extra={
            "city": city,
            "temp": data.get("temperature")
        })
        await ctx.info("Weather data retrieved")
        
        return {
            "success": True,
            "city": city,
            "units": units,
            "data": data,
            "cached": False
        }
    
    # ============================================
    # STEP 3: Handle Specific Errors
    # ============================================
    
    except httpx.TimeoutException as e:
        logger.error("API timeout", extra={
            "city": city,
            "timeout": 10.0
        })
        await ctx.error("Weather service timed out")
        
        return {
            "success": False,
            "error": {
                "code": "API_TIMEOUT",
                "message": "Weather service did not respond in time",
                "details": {
                    "city": city,
                    "timeout_seconds": 10.0
                },
                "retryable": True
            }
        }
    
    except httpx.HTTPStatusError as e:
        status_code = e.response.status_code
        
        if status_code == 404:
            logger.warning(f"City not found", extra={
                "city": city,
                "status": 404
            })
            await ctx.warning(f"City '{city}' not found")
            
            return {
                "success": False,
                "error": {
                    "code": "CITY_NOT_FOUND",
                    "message": f"Weather data not available for '{city}'",
                    "details": {
                        "city": city,
                        "suggestion": "Check city name spelling"
                    },
                    "retryable": False
                }
            }
        
        elif status_code == 429:
            logger.error("Rate limit exceeded", extra={
                "city": city,
                "status": 429
            })
            await ctx.error("Too many requests, please try again later")
            
            return {
                "success": False,
                "error": {
                    "code": "RATE_LIMIT",
                    "message": "Too many requests to weather service",
                    "details": {
                        "retry_after_seconds": 60
                    },
                    "retryable": True
                }
            }
        
        else:
            logger.error(f"HTTP error", extra={
                "city": city,
                "status": status_code,
                "response": e.response.text[:200]
            })
            await ctx.error(f"Weather service error: {status_code}")
            
            return {
                "success": False,
                "error": {
                    "code": "API_ERROR",
                    "message": f"Weather service returned error {status_code}",
                    "details": {"status_code": status_code},
                    "retryable": status_code >= 500
                }
            }
    
    except httpx.NetworkError as e:
        logger.error("Network error", extra={
            "city": city,
            "error": str(e)
        })
        await ctx.error("Network connection failed")
        
        return {
            "success": False,
            "error": {
                "code": "NETWORK_ERROR",
                "message": "Could not connect to weather service",
                "details": {"error": str(e)},
                "retryable": True
            }
        }
    
    except Exception as e:
        # Catch-all for unexpected errors
        logger.exception("Unexpected error", extra={
            "city": city,
            "error_type": type(e).__name__
        })
        await ctx.error("An unexpected error occurred")
        
        return {
            "success": False,
            "error": {
                "code": "INTERNAL_ERROR",
                "message": "An unexpected error occurred",
                "details": {
                    "error_type": type(e).__name__
                },
                "retryable": False
            }
        }
✅ Error Handling Best Practices:

• Always return structured error objects (not just strings)
• Include error codes for programmatic handling
• Mark errors as retryable or not
• Provide actionable error messages
• Log full details server-side, safe messages client-side
• Use retry logic for transient failures
• Never expose internal implementation details in errors

🔖 Versioning: Managing Breaking Changes

As your server evolves, you'll need to update tools without breaking existing clients.

Versioning Strategies

Strategy 1: Semantic Versioning in Server Name

from mcp.server.fastmcp import FastMCP

# Version embedded in server name
mcp = FastMCP("Weather Server v2.1.0")

# Expose version as resource
@mcp.resource("version://info")
def get_version() -> dict:
    return {
        "version": "2.1.0",
        "released": "2024-02-15",
        "breaking_changes": [
            "Removed deprecated 'temp_fahrenheit' field",
            "Changed 'location' parameter to 'city'"
        ],
        "deprecated": [
            "'units' parameter will be required in v3.0.0"
        ]
    }

Strategy 2: Versioned Tool Names

# Support multiple versions simultaneously

@mcp.tool()
async def get_weather_v1(location: str) -> dict:
    """
    DEPRECATED: Use get_weather_v2 instead
    Will be removed in version 3.0.0
    """
    logger.warning("get_weather_v1 called - deprecated")
    # Convert to new format internally
    return await get_weather_v2(city=location)

@mcp.tool()
async def get_weather_v2(city: str, units: str = "metric") -> dict:
    """
    Current version: Get weather data
    
    Args:
        city: City name (replaces 'location' from v1)
        units: Temperature units (new in v2)
    """
    # Implementation...

Strategy 3: Parameter Evolution (Backward Compatible)

from typing import Optional

@mcp.tool()
async def get_weather(
    city: Optional[str] = None,
    location: Optional[str] = None,  # Deprecated, kept for compatibility
    units: str = "metric"
) -> dict:
    """
    Get weather data
    
    Args:
        city: City name (preferred)
        location: DEPRECATED - use 'city' instead
        units: Temperature units
    """
    
    # Handle both old and new parameter names
    if location and not city:
        logger.warning("'location' parameter deprecated, use 'city'")
        city = location
    
    if not city:
        return {
            "success": False,
            "error": {
                "code": "MISSING_PARAMETER",
                "message": "Either 'city' or 'location' is required"
            }
        }
    
    # Rest of implementation...

Deprecation Workflow

Version 2.0.0 (Current) ├─ Feature added: 'units' parameter └─ Old way still works Version 2.1.0 (3 months later) ├─ 'location' parameter marked DEPRECATED ├─ Warnings logged when used └─ Documentation updated Version 2.5.0 (6 months later) ├─ Migration guide published ├─ Warnings become errors in logs └─ "Will be removed in v3.0.0" message Version 3.0.0 (12 months later) ├─ 'location' parameter REMOVED └─ Only 'city' parameter accepted
💡 Versioning Best Practices:

• Give users 6-12 months notice before breaking changes
• Support old + new simultaneously during transition
• Log deprecation warnings prominently
• Provide migration guides and tools
• Use semantic versioning: MAJOR.MINOR.PATCH
• Increment MAJOR for breaking changes only

🎯 Complete Production-Ready Example

Let's put it all together! Here's a production-grade MCP server.

Create production_weather_server.py:

"""
Production-Ready Weather MCP Server
Features:
- Async tools for concurrency
- Comprehensive logging (server + client)
- Structured error handling
- Retry logic with backoff
- Versioning support
- Resource monitoring
"""

import asyncio
import sys
from datetime import datetime
from typing import Optional
import httpx
from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger

# ============================================
# CONFIGURATION
# ============================================

SERVER_VERSION = "2.1.0"
API_TIMEOUT = 10.0
MAX_RETRIES = 3

# ============================================
# LOGGING SETUP
# ============================================

logger = setup_logger("weather_server", log_level="INFO")

# ============================================
# SERVER INITIALIZATION
# ============================================

mcp = FastMCP(f"Weather Server v{SERVER_VERSION}")

logger.info("Server starting", extra={
    "version": SERVER_VERSION,
    "python_version": sys.version,
    "timestamp": datetime.utcnow().isoformat()
})

# ============================================
# UTILITY FUNCTIONS
# ============================================

async def retry_operation(operation, max_attempts: int = MAX_RETRIES):
    """Retry operation with exponential backoff"""
    for attempt in range(max_attempts):
        try:
            return await operation()
        except Exception as e:
            if attempt == max_attempts - 1:
                raise
            delay = (2 ** attempt) + (asyncio.get_event_loop().time() % 1)
            logger.debug(f"Retry {attempt + 1}/{max_attempts}", extra={
                "delay": delay,
                "error": str(e)
            })
            await asyncio.sleep(delay)

def sanitize_city_name(city: str) -> str:
    """Clean and validate city name"""
    return city.strip().title()

# ============================================
# RESOURCE: Server Information
# ============================================

@mcp.resource("server://info")
def server_info() -> dict:
    """Get server version and status information"""
    return {
        "version": SERVER_VERSION,
        "name": "Weather Server",
        "status": "operational",
        "features": [
            "async_tools",
            "retry_logic",
            "structured_errors",
            "client_logging"
        ],
        "api_timeout_seconds": API_TIMEOUT,
        "max_retries": MAX_RETRIES
    }

# ============================================
# TOOL: Get Weather (Current Version)
# ============================================

@mcp.tool()
async def get_weather(
    city: str,
    ctx: Context,
    units: str = "metric"
) -> dict:
    """
    Get current weather for a city
    
    Args:
        city: City name (e.g., "London", "New York")
        units: Temperature units - "metric" (Celsius) or "imperial" (Fahrenheit)
    
    Returns:
        Weather data with temperature, conditions, humidity, etc.
    """
    
    request_id = id(ctx)
    
    logger.info("Tool invoked", extra={
        "tool": "get_weather",
        "city": city,
        "units": units,
        "request_id": request_id
    })
    
    # Validate inputs
    if not city or not city.strip():
        logger.warning("Invalid input: empty city", extra={
            "request_id": request_id
        })
        await ctx.warning("City name cannot be empty")
        return {
            "success": False,
            "error": {
                "code": "INVALID_INPUT",
                "message": "City name is required",
                "retryable": False
            }
        }
    
    city = sanitize_city_name(city)
    
    if units not in ["metric", "imperial"]:
        logger.warning(f"Invalid units: {units}", extra={
            "request_id": request_id
        })
        await ctx.warning(f"Invalid units '{units}', using 'metric'")
        units = "metric"
    
    # Fetch weather data
    await ctx.info(f"Fetching weather for {city}...")
    
    try:
        async def fetch():
            async with httpx.AsyncClient(timeout=API_TIMEOUT) as client:
                logger.debug("API request starting", extra={
                    "city": city,
                    "request_id": request_id
                })
                
                response = await client.get(
                    f"https://api.weather.com/v1/current",
                    params={"city": city, "units": units}
                )
                response.raise_for_status()
                return response.json()
        
        data = await retry_operation(fetch, max_attempts=MAX_RETRIES)
        
        logger.info("Weather retrieved successfully", extra={
            "city": city,
            "temperature": data.get("temperature"),
            "request_id": request_id
        })
        
        await ctx.info("Weather data retrieved")
        
        return {
            "success": True,
            "city": city,
            "units": units,
            "data": data,
            "server_version": SERVER_VERSION,
            "timestamp": datetime.utcnow().isoformat()
        }
    
    except httpx.TimeoutException:
        logger.error("API timeout", extra={
            "city": city,
            "timeout": API_TIMEOUT,
            "request_id": request_id
        })
        await ctx.error(f"Request timed out after {API_TIMEOUT} seconds")
        
        return {
            "success": False,
            "error": {
                "code": "API_TIMEOUT",
                "message": "Weather service did not respond in time",
                "details": {"timeout_seconds": API_TIMEOUT},
                "retryable": True
            }
        }
    
    except httpx.HTTPStatusError as e:
        status = e.response.status_code
        
        logger.error("HTTP error", extra={
            "city": city,
            "status_code": status,
            "request_id": request_id
        })
        
        if status == 404:
            await ctx.warning(f"City '{city}' not found")
            return {
                "success": False,
                "error": {
                    "code": "CITY_NOT_FOUND",
                    "message": f"No weather data for '{city}'",
                    "retryable": False
                }
            }
        
        await ctx.error(f"Weather service error: {status}")
        return {
            "success": False,
            "error": {
                "code": "API_ERROR",
                "message": f"HTTP {status}",
                "retryable": status >= 500
            }
        }
    
    except Exception as e:
        logger.exception("Unexpected error", extra={
            "city": city,
            "request_id": request_id
        })
        await ctx.error("An unexpected error occurred")
        
        return {
            "success": False,
            "error": {
                "code": "INTERNAL_ERROR",
                "message": "Unexpected error",
                "retryable": False
            }
        }

# ============================================
# RUN SERVER
# ============================================

if __name__ == "__main__":
    logger.info("Server ready to accept connections")
    print(f"  Weather Server v{SERVER_VERSION} starting...", file=sys.stderr)
    print(f"Logging to: logs/weather_server.log", file=sys.stderr)
    print(f" Server operational\n", file=sys.stderr)
    
    try:
        mcp.run()
    except KeyboardInterrupt:
        logger.info("Server shutting down (KeyboardInterrupt)")
        print("\n Server shutdown", file=sys.stderr)
    except Exception as e:
        logger.critical("Server crash", extra={
            "error": str(e)
        }, exc_info=True)
        print(f"\n Server crashed: {e}", file=sys.stderr)
        sys.exit(1)

🧪 Exercise: Add Logging and Error Handling

Time to apply what you've learned! Take your existing MCP server and make it production-ready.

Exercise Instructions

🎯 Your Mission:

Take the calculator server from previous lessons and add:
1. Production logging (server + client)
2. Structured error handling
3. Async tool implementation
4. Version information

Starter Code

# Before (basic prototype)
from mcp.server.fastmcp import FastMCP

mcp = FastMCP("Calculator")

@mcp.tool()
def add(a: float, b: float) -> float:
    return a + b

@mcp.tool()
def divide(a: float, b: float) -> float:
    return a / b  # What if b is zero?

if __name__ == "__main__":
    mcp.run()

Solution: Production-Ready Calculator

# After (production-ready)
"""
Production-Ready Calculator MCP Server
"""

from mcp.server.fastmcp import FastMCP, Context
from production_logger import setup_logger
from typing import Union
import sys

# Configuration
VERSION = "1.0.0"

# Setup logging
logger = setup_logger("calculator_server", log_level="INFO")

# Create server
mcp = FastMCP(f"Calculator v{VERSION}")

logger.info("Calculator server initializing", extra={
    "version": VERSION
})

# ============================================
# RESOURCE: Server Info
# ============================================

@mcp.resource("calculator://version")
def get_version() -> dict:
    """Get calculator version information"""
    return {
        "version": VERSION,
        "operations": ["add", "subtract", "multiply", "divide"],
        "status": "operational"
    }

# ============================================
# TOOL: Add Numbers
# ============================================

@mcp.tool()
async def add(a: float, b: float, ctx: Context) -> dict:
    """
    Add two numbers with perfect precision
    
    Args:
        a: First number
        b: Second number
    """
    logger.info("Add operation", extra={
        "a": a,
        "b": b
    })
    
    await ctx.info(f"Adding {a} + {b}...")
    
    try:
        result = a + b
        
        logger.info("Addition successful", extra={
            "result": result
        })
        
        await ctx.info(f"Result: {result}")
        
        return {
            "success": True,
            "operation": "addition",
            "a": a,
            "b": b,
            "result": result,
            "formula": f"{a} + {b} = {result}"
        }
    
    except Exception as e:
        logger.exception("Addition failed", extra={
            "a": a,
            "b": b
        })
        await ctx.error("Addition failed unexpectedly")
        
        return {
            "success": False,
            "error": {
                "code": "CALCULATION_ERROR",
                "message": "Could not perform addition",
                "retryable": False
            }
        }

# ============================================
# TOOL: Divide Numbers (with error handling)
# ============================================

@mcp.tool()
async def divide(a: float, b: float, ctx: Context) -> dict:
    """
    Divide two numbers safely
    
    Args:
        a: Numerator
        b: Denominator (cannot be zero)
    """
    logger.info("Divide operation", extra={
        "a": a,
        "b": b
    })
    
    await ctx.info(f"Dividing {a} ÷ {b}...")
    
    # Handle division by zero
    if b == 0:
        logger.warning("Division by zero attempted", extra={
            "a": a,
            "b": b
        })
        await ctx.warning("Cannot divide by zero")
        
        return {
            "success": False,
            "error": {
                "code": "DIVISION_BY_ZERO",
                "message": "Cannot divide by zero",
                "details": {
                    "numerator": a,
                    "denominator": b
                },
                "retryable": False
            }
        }
    
    try:
        result = a / b
        
        logger.info("Division successful", extra={
            "result": result
        })
        
        await ctx.info(f"Result: {result}")
        
        return {
            "success": True,
            "operation": "division",
            "a": a,
            "b": b,
            "result": result,
            "formula": f"{a} ÷ {b} = {result}"
        }
    
    except Exception as e:
        logger.exception("Division failed", extra={
            "a": a,
            "b": b
        })
        await ctx.error("Division failed unexpectedly")
        
        return {
            "success": False,
            "error": {
                "code": "CALCULATION_ERROR",
                "message": "Could not perform division",
                "retryable": False
            }
        }

# ============================================
# MAIN
# ============================================

if __name__ == "__main__":
    logger.info("Server ready")
    print(f" Calculator Server v{VERSION} starting...", file=sys.stderr)
    print(f" Logging to: logs/calculator_server.log", file=sys.stderr)
    print(f" Server operational\n", file=sys.stderr)
    
    try:
        mcp.run()
    except KeyboardInterrupt:
        logger.info("Server shutdown")
        print("\n Goodbye!", file=sys.stderr)

Testing Your Production Server

# Run server
python production_calculator_server.py

# In another terminal, test with MCP Inspector
mcp dev production_calculator_server.py

# Try these test cases:
# 1. Normal: add(5, 3) → Should succeed
# 2. Error: divide(10, 0) → Should return structured error
# 3. Check logs/calculator_server.log for detailed logging
Verify Your Implementation:

• Async functions (async def)
• Server-side logging (logger.info, logger.error)
• Client-side logging (ctx.info, ctx.error)
• Structured error responses
• Version resource
• Graceful error handling (division by zero)

Production Checklist

Before deploying your MCP server to production, use this checklist:

Async & Performance:
□ All I/O operations are async
□ No blocking calls (time.sleep, requests.get, etc.)
□ Proper connection pooling
□ Timeouts on all external calls
□ Retry logic for transient failures
Logging & Observability:
□ Never print to stdout (STDIO servers)
□ Structured logging format (JSON)
□ Log levels properly used
□ Client-side progress updates (ctx)
□ Sensitive data sanitized from logs
□ Log rotation configured
Error Handling:
□ All errors return structured responses
□ Errors include error codes
□ Retryable flag set correctly
□ User-friendly error messages
□ No internal details leaked
□ Proper exception logging
Versioning & Compatibility:
□ Server version exposed
□ Breaking changes documented
□ Deprecation warnings in place
□ Backward compatibility maintained
□ Migration guides provided
Security & Validation:
□ All inputs validated
□ SQL injection prevention
□ Path traversal protection
□ Rate limiting implemented
□ Authentication/authorization

🚀 Advanced Production Topics

Monitoring & Metrics

import time
from prometheus_client import Counter, Histogram

# Track tool usage
tool_calls = Counter('mcp_tool_calls_total', 'Total tool calls', ['tool_name', 'status'])
tool_duration = Histogram('mcp_tool_duration_seconds', 'Tool execution time', ['tool_name'])

@mcp.tool()
async def monitored_tool(ctx: Context) -> dict:
    start = time.time()
    tool_name = "monitored_tool"
    
    try:
        result = await do_work()
        tool_calls.labels(tool_name=tool_name, status='success').inc()
        return result
    except Exception as e:
        tool_calls.labels(tool_name=tool_name, status='error').inc()
        raise
    finally:
        duration = time.time() - start
        tool_duration.labels(tool_name=tool_name).observe(duration)

Health Checks

@mcp.resource("health://status")
async def health_check() -> dict:
    """Server health status"""
    
    checks = {
        "database": await check_database(),
        "external_api": await check_external_api(),
        "disk_space": check_disk_space()
    }
    
    healthy = all(checks.values())
    
    return {
        "status": "healthy" if healthy else "degraded",
        "checks": checks,
        "timestamp": datetime.utcnow().isoformat()
    }

Graceful Shutdown

import signal
import asyncio

shutdown_event = asyncio.Event()

def handle_shutdown(signum, frame):
    logger.info("Shutdown signal received")
    shutdown_event.set()

signal.signal(signal.SIGTERM, handle_shutdown)
signal.signal(signal.SIGINT, handle_shutdown)

async def cleanup():
    """Cleanup resources before shutdown"""
    logger.info("Performing cleanup...")
    await close_database_connections()
    await flush_logs()
    logger.info("Cleanup complete")

if __name__ == "__main__":
    try:
        mcp.run()
    finally:
        asyncio.run(cleanup())

📚 Summary & Next Steps

What You Learned:

• Async tools prevent blocking and enable concurrency
• Proper logging (stderr, not stdout) enables debugging
• Structured errors help clients handle failures gracefully
• Versioning prevents breaking changes from disrupting users
• Production readiness is about reliability, not just functionality
🎯 Your Path Forward:

Week 1: Convert all your tools to async
Week 2: Add comprehensive logging to every tool
Week 3: Implement structured error handling
Week 4: Add versioning and monitoring
Week 5: Load test and optimize

Essential Resources

Libraries:
• httpx - Async HTTP client
• asyncpg - Async PostgreSQL
• aiofiles - Async file operations
• structlog - Structured logging

Documentation:
• Python AsyncIO: docs.python.org/3/library/asyncio
• MCP Specification: modelcontextprotocol.io
• Logging Best Practices: 12factor.net/logs

✨ Final Thoughts

The difference between a prototype and production isn't complexity—it's reliability.

A prototype works when everything goes right.
A production system works especially when things go wrong.

You've learned:

  • How to handle thousands of concurrent requests efficiently
  • How to debug issues in production without guessing
  • How to fail gracefully and communicate errors clearly
  • How to evolve your server without breaking clients

These aren't "nice to haves"—they're requirements for any system users depend on.

🎯 Your Challenge:

Take your existing MCP server (calculator, weather, or custom).
Spend this week making it production-ready using this guide.

Then deploy it somewhere (local, cloud, anywhere) and use it for real work.

You'll discover edge cases you never imagined.
You'll appreciate good logging when debugging at 2 AM.
You'll thank yourself for structured errors when users report issues.

Production readiness is a journey, not a destination.

Start with async. Add logging. Handle errors. Version carefully.
Monitor everything. Improve iteratively.

Comments