Skip to main content

Building LLM-MCP Agent Loops

Calculating read time…

Imagine having an AI assistant that doesn't just answer questions, but actually does things for you.

It checks your calendar, sends emails, analyzes data, and even writes code — all by deciding which tools to use and when to use them.

This isn't science fiction. It's what happens when you combine Large Language Models (LLMs) with the Model Context Protocol (MCP) in an agent loop.

In this comprehensive guide, you'll learn:

  • How to let LLMs autonomously decide which tools to call
  • How to route LLM outputs to MCP tools
  • How to feed tool results back to the LLM
  • How to build complete agent loops that can handle complex tasks

By the end, you'll build a fully functional AI agent from scratch using only Python.

💡 What You'll Build

A complete research assistant that can search the web, analyze data, and write reports — all by orchestrating multiple tools autonomously.

No frameworks, just pure Python. You'll understand every line.

🎯 Part 1: The Foundation — Understanding Agent Loops

Before we write code, let's understand what an agent loop actually is.

What is an Agent Loop?

An agent loop is a repeating cycle where an AI makes decisions, takes actions, observes results, and repeats.

Think of it like a human solving a problem:

  1. Think: "What should I do next?"
  2. Act: Take an action (use a tool)
  3. Observe: See what happened
  4. Reflect: Did that help? What next?
  5. Repeat: Keep going until the task is complete

This is called the ReAct pattern (Reasoning + Acting).

The Simple Analogy 🧭

Imagine you're cooking a new recipe but don't have all the ingredients.

Your thought process:

  1. "I need eggs" → Check the fridge (tool call)
  2. "No eggs found" → Go to the store (another tool call)
  3. "Bought eggs" → Return home (action complete)
  4. "Now I can cook" → Start cooking (final action)

Each step depends on the previous one. You don't plan everything upfront — you adapt based on what you discover.

That's exactly how an agent loop works.

The Agent Loop Architecture

Here's the complete flow:

┌─────────────────────────────────────────────┐
│           START: User Query                 │
└────────────────┬────────────────────────────┘
                 │
                 ▼
┌─────────────────────────────────────────────┐
│   LLM: Analyze query and decide action     │
│   - Can I answer directly?                  │
│   - Do I need a tool?                       │
│   - Which tool should I use?                │
└────────────────┬────────────────────────────┘
                 │
        ┌────────┴────────┐
        │                 │
        ▼                 ▼
   ┌─────────┐      ┌──────────────┐
   │ Answer  │      │  Tool Call   │
   │ Directly│      │   Request    │
   └────┬────┘      └──────┬───────┘
        │                  │
        │                  ▼
        │         ┌─────────────────┐
        │         │  MCP Server     │
        │         │  Execute Tool   │
        │         └────────┬────────┘
        │                  │
        │                  ▼
        │         ┌─────────────────┐
        │         │  Tool Result    │
        │         └────────┬────────┘
        │                  │
        │                  ▼
        │         ┌─────────────────┐
        │         │ Feed Result     │
        │         │ Back to LLM     │
        │         └────────┬────────┘
        │                  │
        │                  │
        └──────────────────┴─────► LOOP CONTINUES
                                   (until task complete)

The loop continues until the LLM decides it has enough information to answer the user.

🎯 Key Insight

The LLM doesn't execute tools — it generates requests to execute them.

Your code executes the tools and feeds results back to the LLM.

The LLM then decides if it needs more tools or if it can answer the user.

🔧 Part 2: Letting the LLM Decide Which Tool to Call

The first challenge: How does the LLM know which tools are available and when to use them?

Step 1: Define Your Tools

Let's create three simple tools for our agent:

# tools.py
import random
from datetime import datetime

def get_current_time():
    """Get the current date and time."""
    return datetime.now().strftime("%Y-%m-%d %H:%M:%S")

def calculate(expression: str):
    """
    Safely evaluate a mathematical expression.
    
    Args:
        expression: A math expression like "2 + 2" or "15 * 3"
    
    Returns:
        The calculated result
    """
    try:
        # Only allow safe math operations
        allowed_chars = set('0123456789+-*/(). ')
        if not all(c in allowed_chars for c in expression):
            return "Error: Invalid characters in expression"
        
        result = eval(expression)
        return f"{expression} = {result}"
    except Exception as e:
        return f"Error: {str(e)}"

def get_random_fact():
    """Get a random interesting fact."""
    facts = [
        "Honey never spoils. Archaeologists have found 3000-year-old honey that's still edible.",
        "A day on Venus is longer than its year.",
        "Bananas are berries, but strawberries aren't.",
        "There are more stars in the universe than grains of sand on Earth.",
    ]
    return random.choice(facts)

Notice how each function has a clear docstring. This is crucial — the LLM reads these to understand what each tool does.

Step 2: Create Tool Descriptions for the LLM

We need to convert our Python functions into a format the LLM can understand.

# tool_registry.py
import inspect
import json

def get_tool_schema(func):
    """
    Convert a Python function into a tool schema for the LLM.
    """
    sig = inspect.signature(func)
    params = {}
    
    for name, param in sig.parameters.items():
        param_type = "string"  # Default type
        if param.annotation != inspect.Parameter.empty:
            if param.annotation == int:
                param_type = "integer"
            elif param.annotation == float:
                param_type = "number"
            elif param.annotation == bool:
                param_type = "boolean"
        
        params[name] = {
            "type": param_type,
            "description": f"Parameter: {name}"
        }
    
    return {
        "name": func.__name__,
        "description": func.__doc__ or f"Function {func.__name__}",
        "parameters": {
            "type": "object",
            "properties": params,
            "required": list(params.keys())
        }
    }

# Register all available tools
TOOLS = {
    "get_current_time": get_current_time,
    "calculate": calculate,
    "get_random_fact": get_random_fact
}

# Generate schemas
TOOL_SCHEMAS = [get_tool_schema(func) for func in TOOLS.values()]

def get_tools_description():
    """Get a formatted description of all tools for the LLM."""
    descriptions = []
    for schema in TOOL_SCHEMAS:
        params = schema["parameters"]["properties"]
        param_desc = ", ".join([f"{k}: {v['type']}" for k, v in params.items()])
        if param_desc:
            desc = f"{schema['name']}({param_desc}) - {schema['description']}"
        else:
            desc = f"{schema['name']}() - {schema['description']}"
        descriptions.append(desc)
    return "\n".join(descriptions)

This code does something powerful: it automatically generates tool descriptions from your Python functions.

Step 3: Build the LLM Interface

Now we need to communicate with an LLM. We'll use OpenAI's API, but the same pattern works with any LLM.

# llm_client.py
import openai
import os
import json

openai.api_key = os.getenv("OPENAI_API_KEY")

def ask_llm(messages, tools=None):
    """
    Send messages to the LLM and get a response.
    
    Args:
        messages: List of conversation messages
        tools: Optional list of tool schemas
    
    Returns:
        The LLM's response
    """
    params = {
        "model": "gpt-4",
        "messages": messages,
        "temperature": 0.7,
    }
    
    if tools:
        params["tools"] = [{"type": "function", "function": tool} for tool in tools]
        params["tool_choice"] = "auto"  # Let LLM decide
    
    response = openai.chat.completions.create(**params)
    return response.choices[0].message

def parse_tool_calls(message):
    """
    Extract tool calls from the LLM's response.
    
    Returns:
        List of tool calls with name and arguments
    """
    if not message.tool_calls:
        return []
    
    tool_calls = []
    for call in message.tool_calls:
        tool_calls.append({
            "id": call.id,
            "name": call.function.name,
            "arguments": json.loads(call.function.arguments)
        })
    
    return tool_calls

Here's what this does:

  • Sends your conversation to the LLM
  • Provides tool schemas so the LLM knows what's available
  • Lets the LLM decide whether to use a tool (tool_choice: "auto")
  • Parses any tool calls from the response

✅ How LLM Tool Selection Works

The LLM reads your tool descriptions and compares them to the user's query.

If the user asks "What time is it?" the LLM matches this to get_current_time()

If the user asks "What's 234 * 56?" the LLM matches this to calculate(expression)

The LLM generates a JSON tool call request, not the actual execution.

🔄 Part 3: Routing LLM Output to MCP Tools

Now comes the exciting part: actually executing the tools the LLM requested.

Understanding the Routing Process

When the LLM decides to use a tool, it returns something like this:

{
  "id": "call_abc123",
  "name": "calculate",
  "arguments": {
    "expression": "234 * 56"
  }
}

Your job is to:

  1. Read this JSON
  2. Find the corresponding tool function
  3. Call it with the provided arguments
  4. Capture the result

Building the Tool Executor

# tool_executor.py
from tool_registry import TOOLS
import json

def execute_tool(tool_name, arguments):
    """
    Execute a tool by name with given arguments.
    
    Args:
        tool_name: Name of the tool to execute
        arguments: Dictionary of arguments for the tool
    
    Returns:
        The tool's result as a string
    """
    if tool_name not in TOOLS:
        return f"Error: Tool '{tool_name}' not found"
    
    try:
        tool_function = TOOLS[tool_name]
        
        # Call the function with unpacked arguments
        result = tool_function(**arguments)
        
        return str(result)
    
    except TypeError as e:
        return f"Error: Invalid arguments for {tool_name}: {str(e)}"
    except Exception as e:
        return f"Error executing {tool_name}: {str(e)}"

def execute_tool_calls(tool_calls):
    """
    Execute multiple tool calls and return results.
    
    Args:
        tool_calls: List of tool call objects
    
    Returns:
        List of results with tool call IDs
    """
    results = []
    
    for call in tool_calls:
        tool_id = call["id"]
        tool_name = call["name"]
        arguments = call["arguments"]
        
        print(f"\n🔧 Executing tool: {tool_name}")
        print(f"   Arguments: {json.dumps(arguments, indent=2)}")
        
        result = execute_tool(tool_name, arguments)
        
        print(f"   Result: {result}\n")
        
        results.append({
            "tool_call_id": tool_id,
            "role": "tool",
            "name": tool_name,
            "content": result
        })
    
    return results

This is your router — it takes LLM output and executes the right tools.

Example Flow

Let's trace what happens step by step:

User asks: "What's 15 times 23?"

  1. LLM receives: The question + list of available tools
  2. LLM thinks: "This is a math question. I should use the calculate tool."
  3. LLM returns:
    {
      "name": "calculate",
      "arguments": {"expression": "15 * 23"}
    }
    
  4. Your code executes: calculate("15 * 23")
  5. Tool returns: "15 * 23 = 345"

But we're not done! We need to feed this result back to the LLM.

🔁 Part 4: Feeding Tool Results Back to the LLM

After executing a tool, we need to give the result back to the LLM so it can continue reasoning.

The Feedback Loop

Here's how it works:

  1. User asks a question
  2. LLM requests a tool
  3. You execute the tool
  4. You append the result to the conversation
  5. You send the updated conversation back to the LLM
  6. LLM reads the result and decides what to do next

The key is maintaining a conversation history that includes:

  • User messages
  • LLM responses
  • Tool calls
  • Tool results

Building the Conversation Manager

# conversation.py

class Conversation:
    """Manages the conversation history between user, LLM, and tools."""
    
    def __init__(self, system_prompt):
        self.messages = [
            {"role": "system", "content": system_prompt}
        ]
    
    def add_user_message(self, content):
        """Add a user message to the conversation."""
        self.messages.append({
            "role": "user",
            "content": content
        })
    
    def add_assistant_message(self, message):
        """Add an LLM response to the conversation."""
        msg = {
            "role": "assistant",
            "content": message.content or ""
        }
        
        # Include tool calls if present
        if message.tool_calls:
            msg["tool_calls"] = [
                {
                    "id": call.id,
                    "type": "function",
                    "function": {
                        "name": call.function.name,
                        "arguments": call.function.arguments
                    }
                }
                for call in message.tool_calls
            ]
        
        self.messages.append(msg)
    
    def add_tool_results(self, results):
        """Add tool execution results to the conversation."""
        for result in results:
            self.messages.append(result)
    
    def get_messages(self):
        """Get the complete conversation history."""
        return self.messages
    
    def print_conversation(self):
        """Print the conversation in a readable format."""
        for msg in self.messages:
            role = msg["role"].upper()
            content = msg.get("content", "")
            
            if role == "SYSTEM":
                print(f"\n{'='*50}")
                print(f"SYSTEM: {content}")
                print(f"{'='*50}\n")
            elif role == "USER":
                print(f"\n👤 USER: {content}\n")
            elif role == "ASSISTANT":
                if content:
                    print(f"🤖 ASSISTANT: {content}\n")
                if "tool_calls" in msg:
                    print("   [Requesting tool calls...]")
            elif role == "TOOL":
                print(f"🔧 TOOL ({msg['name']}): {content}\n")

This class keeps track of everything that happens in the conversation.

💡 Why Conversation History Matters

The LLM is stateless — it doesn't remember previous interactions.

By sending the complete conversation history each time, the LLM has context.

This allows it to make informed decisions based on what tools were called and what results came back.

♾️ Part 5: Building Complete Agent Loops

Now we combine everything into a full agent loop that can handle complex multi-step tasks.

The Agent Loop Implementation

# agent.py
from llm_client import ask_llm, parse_tool_calls
from tool_executor import execute_tool_calls
from tool_registry import TOOL_SCHEMAS, get_tools_description
from conversation import Conversation

class Agent:
    """
    An autonomous agent that uses LLM + MCP tools to solve tasks.
    """
    
    def __init__(self, max_iterations=10):
        """
        Initialize the agent.
        
        Args:
            max_iterations: Maximum number of loop iterations to prevent infinite loops
        """
        self.max_iterations = max_iterations
        
        # Create system prompt with tool descriptions
        tools_desc = get_tools_description()
        self.system_prompt = f"""You are a helpful AI assistant with access to tools.

Available tools:
{tools_desc}

When you need to use a tool, request it and wait for the result.
When you have enough information, provide a final answer to the user.
Be concise and helpful."""
    
    def run(self, user_query):
        """
        Run the agent loop for a user query.
        
        Args:
            user_query: The user's question or task
        
        Returns:
            The final answer
        """
        print(f"\n{'='*60}")
        print(f"🚀 Starting Agent Loop")
        print(f"{'='*60}\n")
        print(f"👤 User Query: {user_query}\n")
        
        # Initialize conversation
        conversation = Conversation(self.system_prompt)
        conversation.add_user_message(user_query)
        
        iteration = 0
        
        # Main agent loop
        while iteration < self.max_iterations:
            iteration += 1
            print(f"\n--- Iteration {iteration} ---\n")
            
            # Get LLM response
            llm_response = ask_llm(
                conversation.get_messages(),
                tools=TOOL_SCHEMAS
            )
            
            # Add LLM response to conversation
            conversation.add_assistant_message(llm_response)
            
            # Check if LLM wants to use tools
            tool_calls = parse_tool_calls(llm_response)
            
            if not tool_calls:
                # No tool calls - LLM has final answer
                final_answer = llm_response.content
                print(f"\n{'='*60}")
                print(f"✅ Final Answer:")
                print(f"{'='*60}\n")
                print(f"{final_answer}\n")
                
                return final_answer
            
            # Execute tools
            print(f"🔧 Executing {len(tool_calls)} tool(s)...\n")
            tool_results = execute_tool_calls(tool_calls)
            
            # Feed results back to conversation
            conversation.add_tool_results(tool_results)
            
            # Loop continues...
        
        # Max iterations reached
        print("\n⚠️  Max iterations reached. Stopping loop.\n")
        return "I couldn't complete the task within the iteration limit."

# Example usage
if __name__ == "__main__":
    agent = Agent(max_iterations=5)
    
    # Example 1: Simple calculation
    agent.run("What's 456 multiplied by 789?")
    
    # Example 2: Multiple tools
    agent.run("What time is it, and can you also tell me a random fact?")

This is a complete, working agent! Let's break down what happens:

  1. User provides a query
  2. Agent sends query + tool schemas to LLM
  3. LLM decides if it needs tools
  4. If yes → Agent executes tools and feeds results back (loop continues)
  5. If no → LLM provides final answer (loop ends)

Example Agent Execution

Let's trace a real example: "What's 234 * 56 and what time is it?"

Iteration 1:

LLM thinks: "User wants calculation AND current time. I need two tools."

LLM requests:
  1. calculate(expression="234 * 56")
  2. get_current_time()

Agent executes both tools:
  - calculate returns: "234 * 56 = 13104"
  - get_current_time returns: "2026-02-04 14:30:25"

Agent feeds results back to LLM...

Iteration 2:

LLM reads tool results.

LLM thinks: "I now have both pieces of information. I can answer."

LLM responds:
  "The calculation 234 * 56 equals 13,104. 
   The current time is 2:30 PM on February 4th, 2026."

No tool calls → Loop ends with final answer.

The agent handled a multi-step task completely autonomously!

✅ Agent Loop Best Practices

1. Set max iterations to prevent infinite loops
2. Log everything for debugging
3. Handle errors gracefully in tool execution
4. Give clear tool descriptions to guide LLM
5. Include system prompts that explain the agent's role

🎯 Part 6: Advanced Agent Patterns

Now that you understand the basics, let's explore advanced patterns.

Pattern 1: Conditional Tool Calling

Sometimes you want the LLM to always use a specific tool first.

# Force tool usage on first call
llm_response = ask_llm(
    messages,
    tools=TOOL_SCHEMAS,
    tool_choice={"type": "function", "function": {"name": "search_database"}}
)

This forces the LLM to call search_database first, then decide what to do with the results.

Pattern 2: Tool Chaining

Build tools that depend on other tools' outputs.

def search_products(query: str):
    """Search for products matching a query."""
    # Returns list of product IDs
    return ["prod_123", "prod_456", "prod_789"]

def get_product_details(product_id: str):
    """Get detailed information about a specific product."""
    # Returns product details
    return {
        "id": product_id,
        "name": "Example Product",
        "price": 99.99
    }

The LLM can chain these:

  1. Call search_products("laptop") → Get IDs
  2. Call get_product_details("prod_123") → Get details
  3. Analyze details and answer user

Pattern 3: Parallel Tool Execution

When tools don't depend on each other, execute them in parallel for speed.

import asyncio

async def execute_tool_async(tool_name, arguments):
    """Execute a single tool asynchronously."""
    # Your tool execution logic
    return execute_tool(tool_name, arguments)

async def execute_tools_parallel(tool_calls):
    """Execute multiple tools in parallel."""
    tasks = [
        execute_tool_async(call["name"], call["arguments"])
        for call in tool_calls
    ]
    results = await asyncio.gather(*tasks)
    return results

If the LLM requests weather for 3 different cities, run all 3 API calls simultaneously.

Pattern 4: Memory and State Management

For long-running agents, add memory storage.

class StatefulAgent(Agent):
    """Agent with persistent memory."""
    
    def __init__(self):
        super().__init__()
        self.memory = {}  # Store information across runs
    
    def store_memory(self, key, value):
        """Store information for future reference."""
        self.memory[key] = value
    
    def recall_memory(self, key):
        """Recall previously stored information."""
        return self.memory.get(key)

Now tools can save and retrieve information across multiple user queries.

🏗️ Part 7: Building a Real-World Research Assistant

Let's build a complete research assistant that can search, analyze, and report.

Define Research Tools

# research_tools.py
import requests
from bs4 import BeautifulSoup

def web_search(query: str, num_results: int = 3):
    """
    Search the web and return top results.
    
    Args:
        query: Search query
        num_results: Number of results to return
    
    Returns:
        List of search results with titles and URLs
    """
    # In production, use a real search API like Google Custom Search
    # This is a simplified example
    return [
        {
            "title": f"Result {i+1} for '{query}'",
            "url": f"https://example.com/result{i+1}",
            "snippet": f"Information about {query}..."
        }
        for i in range(num_results)
    ]

def fetch_webpage(url: str):
    """
    Fetch and extract text content from a webpage.
    
    Args:
        url: The URL to fetch
    
    Returns:
        Extracted text content
    """
    try:
        response = requests.get(url, timeout=10)
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # Extract text from paragraphs
        paragraphs = soup.find_all('p')
        text = ' '.join([p.get_text() for p in paragraphs[:10]])  # First 10 paragraphs
        
        return text[:1000]  # Limit to 1000 characters
    except Exception as e:
        return f"Error fetching webpage: {str(e)}"

def analyze_data(data: str, analysis_type: str):
    """
    Analyze text data and extract insights.
    
    Args:
        data: The text data to analyze
        analysis_type: Type of analysis ('summary', 'keywords', 'sentiment')
    
    Returns:
        Analysis results
    """
    if analysis_type == "summary":
        # Simplified summary - in production use NLP models
        words = data.split()
        return f"Data contains {len(words)} words. " + " ".join(words[:50]) + "..."
    
    elif analysis_type == "keywords":
        # Extract most common words
        words = data.lower().split()
        word_freq = {}
        for word in words:
            if len(word) > 4:  # Only words longer than 4 chars
                word_freq[word] = word_freq.get(word, 0) + 1
        
        top_words = sorted(word_freq.items(), key=lambda x: x[1], reverse=True)[:5]
        return f"Top keywords: {', '.join([w[0] for w in top_words])}"
    
    elif analysis_type == "sentiment":
        # Simplified sentiment - in production use sentiment analysis models
        positive_words = ["good", "great", "excellent", "amazing", "wonderful"]
        negative_words = ["bad", "terrible", "awful", "poor", "horrible"]
        
        text_lower = data.lower()
        pos_count = sum(text_lower.count(word) for word in positive_words)
        neg_count = sum(text_lower.count(word) for word in negative_words)
        
        if pos_count > neg_count:
            return f"Sentiment: Positive (pos: {pos_count}, neg: {neg_count})"
        elif neg_count > pos_count:
            return f"Sentiment: Negative (pos: {pos_count}, neg: {neg_count})"
        else:
            return f"Sentiment: Neutral (pos: {pos_count}, neg: {neg_count})"
    
    return "Unknown analysis type"

def save_report(filename: str, content: str):
    """
    Save research findings to a file.
    
    Args:
        filename: Name of the file
        content: Content to save
    
    Returns:
        Confirmation message
    """
    try:
        with open(filename, 'w') as f:
            f.write(content)
        return f"Report saved to {filename}"
    except Exception as e:
        return f"Error saving report: {str(e)}"

Create the Research Agent

# research_agent.py
from agent import Agent
from tool_registry import TOOLS
from research_tools import web_search, fetch_webpage, analyze_data, save_report

# Register research tools
TOOLS.update({
    "web_search": web_search,
    "fetch_webpage": fetch_webpage,
    "analyze_data": analyze_data,
    "save_report": save_report
})

class ResearchAgent(Agent):
    """
    Specialized agent for research tasks.
    """
    
    def __init__(self):
        super().__init__(max_iterations=15)  # More iterations for complex research
        
        # Enhanced system prompt for research
        tools_desc = self.get_tools_description()
        self.system_prompt = f"""You are an expert research assistant.

Your capabilities:
1. Search the web for information
2. Fetch and read webpage content
3. Analyze data for insights
4. Save research reports

Available tools:
{tools_desc}

Research Process:
1. Search for relevant information
2. Fetch detailed content from promising sources
3. Analyze the gathered data
4. Synthesize findings into a coherent report
5. Save the report if requested

Be thorough, accurate, and cite your sources."""

# Example usage
if __name__ == "__main__":
    agent = ResearchAgent()
    
    result = agent.run(
        "Research the benefits of Python for data science and save a summary report to python_research.txt"
    )

Example Research Session

Let's trace what happens when you ask: "Research AI safety and summarize the top 3 concerns"

Iteration 1: LLM calls web_search("AI safety concerns")

Iteration 2: LLM calls fetch_webpage(url1), fetch_webpage(url2), fetch_webpage(url3)

Iteration 3: LLM calls analyze_data(combined_text, "summary")

Iteration 4: LLM synthesizes all information and provides final answer with top 3 concerns

The agent orchestrated multiple tools across multiple steps completely autonomously!

🐛 Part 8: Error Handling and Recovery

Real-world agents need robust error handling.

Common Failure Modes

  1. Tool execution fails (API down, invalid input)
  2. LLM hallucinations (requests non-existent tools)
  3. Infinite loops (LLM keeps calling same tool)
  4. Timeout errors (slow APIs, network issues)

Enhanced Error Handling

# robust_agent.py
import time
from collections import Counter

class RobustAgent(Agent):
    """Agent with comprehensive error handling."""
    
    def __init__(self, max_iterations=10, timeout_seconds=60):
        super().__init__(max_iterations)
        self.timeout_seconds = timeout_seconds
        self.start_time = None
        self.tool_call_history = []
    
    def check_timeout(self):
        """Check if agent has exceeded time limit."""
        elapsed = time.time() - self.start_time
        if elapsed > self.timeout_seconds:
            raise TimeoutError(f"Agent exceeded {self.timeout_seconds}s time limit")
    
    def detect_loop(self):
        """Detect if agent is stuck in a loop."""
        if len(self.tool_call_history) < 3:
            return False
        
        # Check last 3 calls
        recent_calls = self.tool_call_history[-3:]
        counter = Counter(recent_calls)
        
        # If same tool called 3 times in a row, we're in a loop
        if counter.most_common(1)[0][1] >= 3:
            return True
        
        return False
    
    def execute_tool_safely(self, tool_name, arguments):
        """Execute tool with error handling and retry logic."""
        max_retries = 3
        retry_delay = 1  # seconds
        
        for attempt in range(max_retries):
            try:
                result = execute_tool(tool_name, arguments)
                
                # Track successful call
                self.tool_call_history.append(tool_name)
                
                return result
            
            except Exception as e:
                if attempt < max_retries - 1:
                    print(f"⚠️  Tool failed (attempt {attempt + 1}/{max_retries}): {str(e)}")
                    print(f"   Retrying in {retry_delay}s...")
                    time.sleep(retry_delay)
                    retry_delay *= 2  # Exponential backoff
                else:
                    error_msg = f"Tool {tool_name} failed after {max_retries} attempts: {str(e)}"
                    print(f"❌ {error_msg}")
                    return error_msg
    
    def run(self, user_query):
        """Run agent with comprehensive error handling."""
        self.start_time = time.time()
        self.tool_call_history = []
        
        try:
            iteration = 0
            conversation = Conversation(self.system_prompt)
            conversation.add_user_message(user_query)
            
            while iteration < self.max_iterations:
                iteration += 1
                
                # Check timeout
                self.check_timeout()
                
                # Check for loops
                if self.detect_loop():
                    print("\n⚠️  Loop detected! Stopping to prevent infinite recursion.\n")
                    return "I seem to be stuck in a loop. Let me try a different approach."
                
                # Get LLM response with error handling
                try:
                    llm_response = ask_llm(conversation.get_messages(), tools=TOOL_SCHEMAS)
                except Exception as e:
                    print(f"\n❌ LLM error: {str(e)}\n")
                    return f"I encountered an error while processing: {str(e)}"
                
                conversation.add_assistant_message(llm_response)
                tool_calls = parse_tool_calls(llm_response)
                
                if not tool_calls:
                    return llm_response.content
                
                # Execute tools safely
                tool_results = []
                for call in tool_calls:
                    result = self.execute_tool_safely(call["name"], call["arguments"])
                    tool_results.append({
                        "tool_call_id": call["id"],
                        "role": "tool",
                        "name": call["name"],
                        "content": result
                    })
                
                conversation.add_tool_results(tool_results)
            
            return "Max iterations reached without completing the task."
        
        except TimeoutError as e:
            return f"Operation timed out: {str(e)}"
        except Exception as e:
            return f"Unexpected error: {str(e)}"

This enhanced agent handles:

  • Timeouts (prevents hanging)
  • Infinite loops (detects repeated tool calls)
  • Tool failures (retries with exponential backoff)
  • LLM errors (graceful degradation)

❌ Common Pitfalls to Avoid

1. No max iterations: Agent runs forever
2. No error handling: One failed tool crashes everything
3. Vague tool descriptions: LLM doesn't know when to use them
4. Not logging: Impossible to debug what went wrong
5. Synchronous execution: Slow for parallel operations

📊 Part 9: Monitoring and Debugging

Production agents need monitoring to understand what's happening.

Adding Comprehensive Logging

# logger.py
import json
from datetime import datetime

class AgentLogger:
    """Logger for tracking agent execution."""
    
    def __init__(self, log_file="agent_log.json"):
        self.log_file = log_file
        self.current_run = {
            "start_time": None,
            "end_time": None,
            "user_query": None,
            "iterations": [],
            "final_answer": None,
            "errors": []
        }
    
    def start_run(self, user_query):
        """Start logging a new agent run."""
        self.current_run = {
            "start_time": datetime.now().isoformat(),
            "user_query": user_query,
            "iterations": [],
            "errors": []
        }
    
    def log_iteration(self, iteration_num, tool_calls, tool_results):
        """Log an iteration."""
        self.current_run["iterations"].append({
            "iteration": iteration_num,
            "timestamp": datetime.now().isoformat(),
            "tool_calls": tool_calls,
            "results": tool_results
        })
    
    def log_error(self, error):
        """Log an error."""
        self.current_run["errors"].append({
            "timestamp": datetime.now().isoformat(),
            "error": str(error)
        })
    
    def end_run(self, final_answer):
        """Complete the run log."""
        self.current_run["end_time"] = datetime.now().isoformat()
        self.current_run["final_answer"] = final_answer
        
        # Save to file
        with open(self.log_file, 'a') as f:
            f.write(json.dumps(self.current_run, indent=2) + "\n")
    
    def get_stats(self):
        """Get statistics about the run."""
        return {
            "total_iterations": len(self.current_run["iterations"]),
            "total_tool_calls": sum(
                len(it["tool_calls"]) for it in self.current_run["iterations"]
            ),
            "errors": len(self.current_run["errors"]),
            "duration": self._calculate_duration()
        }
    
    def _calculate_duration(self):
        """Calculate run duration."""
        if not self.current_run["start_time"] or not self.current_run["end_time"]:
            return None
        
        start = datetime.fromisoformat(self.current_run["start_time"])
        end = datetime.fromisoformat(self.current_run["end_time"])
        return (end - start).total_seconds()

Visualization Dashboard

Create a simple dashboard to monitor agent performance:

# dashboard.py
import json
from collections import Counter

def analyze_logs(log_file="agent_log.json"):
    """Analyze agent logs and generate insights."""
    
    with open(log_file, 'r') as f:
        content = f.read()
        runs = [json.loads(line) for line in content.strip().split('\n') if line]
    
    total_runs = len(runs)
    successful_runs = sum(1 for r in runs if r.get("final_answer") and not r.get("errors"))
    
    all_tool_calls = []
    for run in runs:
        for iteration in run.get("iterations", []):
            for call in iteration.get("tool_calls", []):
                all_tool_calls.append(call["name"])
    
    tool_usage = Counter(all_tool_calls)
    
    avg_iterations = sum(len(r.get("iterations", [])) for r in runs) / total_runs if total_runs > 0 else 0
    
    print("\n" + "="*60)
    print("AGENT PERFORMANCE DASHBOARD")
    print("="*60 + "\n")
    print(f"Total Runs: {total_runs}")
    print(f"Successful: {successful_runs} ({successful_runs/total_runs*100:.1f}%)")
    print(f"Failed: {total_runs - successful_runs}")
    print(f"\nAverage Iterations per Run: {avg_iterations:.1f}")
    print(f"\nMost Used Tools:")
    for tool, count in tool_usage.most_common(5):
        print(f"  {tool}: {count} calls")
    print("\n" + "="*60 + "\n")

# Run analysis
if __name__ == "__main__":
    analyze_logs()

🎓 Part 10: Best Practices and Optimization

Let's wrap up with production-ready best practices.

1. Optimize Token Usage

LLM API costs are based on tokens. Reduce costs by:

  • Truncating long tool results (first 500 chars)
  • Summarizing web content before feeding to LLM
  • Using cheaper models for simple tool selection
  • Caching tool results that don't change

2. Implement Caching

import hashlib
import json

class ToolCache:
    """Cache tool results to avoid redundant calls."""
    
    def __init__(self):
        self.cache = {}
    
    def _get_key(self, tool_name, arguments):
        """Generate cache key from tool and args."""
        key_str = f"{tool_name}:{json.dumps(arguments, sort_keys=True)}"
        return hashlib.md5(key_str.encode()).hexdigest()
    
    def get(self, tool_name, arguments):
        """Get cached result if available."""
        key = self._get_key(tool_name, arguments)
        return self.cache.get(key)
    
    def set(self, tool_name, arguments, result):
        """Cache a tool result."""
        key = self._get_key(tool_name, arguments)
        self.cache[key] = result

# Use in agent
cache = ToolCache()

def execute_tool_with_cache(tool_name, arguments):
    """Execute tool with caching."""
    cached = cache.get(tool_name, arguments)
    if cached:
        print(f"✅ Using cached result for {tool_name}")
        return cached
    
    result = execute_tool(tool_name, arguments)
    cache.set(tool_name, arguments, result)
    return result

3. Rate Limiting

import time
from collections import deque

class RateLimiter:
    """Rate limiter for API calls."""
    
    def __init__(self, max_calls, time_window):
        self.max_calls = max_calls
        self.time_window = time_window  # seconds
        self.calls = deque()
    
    def acquire(self):
        """Wait until rate limit allows next call."""
        now = time.time()
        
        # Remove old calls outside time window
        while self.calls and self.calls[0] < now - self.time_window:
            self.calls.popleft()
        
        # Wait if at limit
        if len(self.calls) >= self.max_calls:
            sleep_time = self.calls[0] + self.time_window - now
            if sleep_time > 0:
                time.sleep(sleep_time)
                self.acquire()  # Retry
        
        self.calls.append(now)

# 10 calls per minute
rate_limiter = RateLimiter(max_calls=10, time_window=60)

def rate_limited_llm_call(messages, tools):
    """Make LLM call with rate limiting."""
    rate_limiter.acquire()
    return ask_llm(messages, tools)

4. Testing Your Agent

# test_agent.py
import pytest
from agent import Agent

def test_simple_calculation():
    """Test agent can handle basic math."""
    agent = Agent()
    result = agent.run("What's 5 + 7?")
    assert "12" in result

def test_multiple_tools():
    """Test agent can use multiple tools."""
    agent = Agent()
    result = agent.run("What time is it and what's 10 * 10?")
    assert "100" in result  # Calculation result
    # Time check would need mocking

def test_error_handling():
    """Test agent handles tool errors gracefully."""
    agent = Agent()
    result = agent.run("Calculate 'invalid expression'")
    # Should return error message, not crash
    assert result is not None

def test_max_iterations():
    """Test agent stops at max iterations."""
    agent = Agent(max_iterations=2)
    # Complex query that would need many iterations
    result = agent.run("Solve world hunger")
    # Should stop after 2 iterations
    assert "iteration limit" in result.lower() or result is not None

if __name__ == "__main__":
    pytest.main([__file__])

🚀 Putting It All Together: Complete Production Agent

Here's a complete, production-ready agent with all best practices:

# production_agent.py
import time
from typing import Optional
from llm_client import ask_llm, parse_tool_calls
from tool_executor import execute_tool
from conversation import Conversation
from logger import AgentLogger
from tool_cache import ToolCache
from rate_limiter import RateLimiter

class ProductionAgent:
    """
    Production-ready agent with full error handling, logging, caching, and monitoring.
    """
    
    def __init__(
        self,
        tools,
        tool_schemas,
        max_iterations=10,
        timeout_seconds=120,
        enable_cache=True,
        rate_limit_calls=60  # calls per minute
    ):
        self.tools = tools
        self.tool_schemas = tool_schemas
        self.max_iterations = max_iterations
        self.timeout_seconds = timeout_seconds
        
        # Initialize components
        self.logger = AgentLogger()
        self.cache = ToolCache() if enable_cache else None
        self.rate_limiter = RateLimiter(max_calls=rate_limit_calls, time_window=60)
        
        # State tracking
        self.tool_call_history = []
        self.start_time = None
    
    def run(self, user_query: str) -> str:
        """
        Execute the agent loop with full production features.
        """
        self.start_time = time.time()
        self.logger.start_run(user_query)
        
        print(f"\n{'='*60}")
        print(f"🚀 Production Agent Starting")
        print(f"{'='*60}\n")
        print(f"Query: {user_query}\n")
        
        try:
            conversation = Conversation(self._get_system_prompt())
            conversation.add_user_message(user_query)
            
            iteration = 0
            
            while iteration < self.max_iterations:
                iteration += 1
                print(f"\n--- Iteration {iteration} ---")
                
                # Timeout check
                if time.time() - self.start_time > self.timeout_seconds:
                    raise TimeoutError("Agent exceeded time limit")
                
                # Loop detection
                if self._detect_loop():
                    raise RuntimeError("Infinite loop detected")
                
                # Get LLM response with rate limiting
                self.rate_limiter.acquire()
                llm_response = ask_llm(
                    conversation.get_messages(),
                    tools=self.tool_schemas
                )
                
                conversation.add_assistant_message(llm_response)
                tool_calls = parse_tool_calls(llm_response)
                
                # No tools needed - we have final answer
                if not tool_calls:
                    final_answer = llm_response.content
                    self.logger.end_run(final_answer)
                    self._print_stats()
                    return final_answer
                
                # Execute tools with caching and error handling
                tool_results = self._execute_tools(tool_calls)
                conversation.add_tool_results(tool_results)
                
                # Log iteration
                self.logger.log_iteration(iteration, tool_calls, tool_results)
            
            # Max iterations reached
            result = "Task incomplete: maximum iterations reached"
            self.logger.end_run(result)
            self._print_stats()
            return result
        
        except Exception as e:
            self.logger.log_error(e)
            error_msg = f"Error: {str(e)}"
            self.logger.end_run(error_msg)
            print(f"\n❌ {error_msg}\n")
            return error_msg
    
    def _execute_tools(self, tool_calls):
        """Execute tools with caching and retry logic."""
        results = []
        
        for call in tool_calls:
            tool_name = call["name"]
            arguments = call["arguments"]
            
            # Try cache first
            if self.cache:
                cached = self.cache.get(tool_name, arguments)
                if cached:
                    print(f"✅ Cache hit: {tool_name}")
                    results.append({
                        "tool_call_id": call["id"],
                        "role": "tool",
                        "name": tool_name,
                        "content": cached
                    })
                    continue
            
            # Execute with retry
            result = self._execute_with_retry(tool_name, arguments)
            
            # Cache result
            if self.cache:
                self.cache.set(tool_name, arguments, result)
            
            # Track for loop detection
            self.tool_call_history.append(tool_name)
            
            results.append({
                "tool_call_id": call["id"],
                "role": "tool",
                "name": tool_name,
                "content": result
            })
        
        return results
    
    def _execute_with_retry(self, tool_name, arguments, max_retries=3):
        """Execute tool with exponential backoff retry."""
        retry_delay = 1
        
        for attempt in range(max_retries):
            try:
                return execute_tool(tool_name, arguments)
            except Exception as e:
                if attempt < max_retries - 1:
                    print(f"⚠️  Retry {attempt + 1}/{max_retries} for {tool_name}")
                    time.sleep(retry_delay)
                    retry_delay *= 2
                else:
                    return f"Error after {max_retries} attempts: {str(e)}"
    
    def _detect_loop(self):
        """Detect if agent is stuck in a loop."""
        if len(self.tool_call_history) < 4:
            return False
        
        recent = self.tool_call_history[-4:]
        return len(set(recent)) == 1  # Same tool 4 times
    
    def _get_system_prompt(self):
        """Generate system prompt with tool descriptions."""
        tools_desc = "\n".join([
            f"- {s['name']}: {s['description']}"
            for s in self.tool_schemas
        ])
        
        return f"""You are a helpful AI assistant with access to tools.

Available tools:
{tools_desc}

Use tools when needed to accomplish tasks. When you have sufficient information, provide a final answer."""
    
    def _print_stats(self):
        """Print execution statistics."""
        stats = self.logger.get_stats()
        print(f"\n{'='*60}")
        print(f"📊 Execution Stats")
        print(f"{'='*60}")
        print(f"Iterations: {stats['total_iterations']}")
        print(f"Tool calls: {stats['total_tool_calls']}")
        print(f"Errors: {stats['errors']}")
        print(f"Duration: {stats['duration']:.2f}s")
        print(f"{'='*60}\n")

This production agent includes:

  • ✅ Comprehensive error handling
  • ✅ Tool result caching
  • ✅ Rate limiting
  • ✅ Timeout protection
  • ✅ Loop detection
  • ✅ Retry logic
  • ✅ Complete logging
  • ✅ Performance monitoring

Let's recap your complete journey:

Part 1: Understanding Agent Loops

  • Agent loops follow the ReAct pattern (Reason → Act → Observe)
  • LLMs decide actions, your code executes them
  • Loop continues until task completion

Part 2: LLM Tool Selection

  • Define tools with clear docstrings
  • Convert functions to schemas the LLM can read
  • Let LLM decide which tool to use based on user query

Part 3: Routing to MCP Tools

  • Parse LLM's tool call requests (JSON format)
  • Execute the requested tool with provided arguments
  • Capture results for feedback

Part 4: Feedback Loop

  • Maintain complete conversation history
  • Append tool results to conversation
  • Send updated history back to LLM
  • LLM uses results to continue reasoning

Part 5: Production Features

  • Error handling and retries
  • Caching for efficiency
  • Rate limiting for API protection
  • Logging and monitoring
  • Loop and timeout detection

📖 Additional Resources

  • MCP Protocol Specification: modelcontextprotocol.io
  • OpenAI Function Calling: platform.openai.com/docs/guides/function-calling
  • Anthropic Claude Tool Use: docs.anthropic.com/claude/docs/tool-use
  • ReAct Paper: Reason and Act pattern research
  • LangChain Agents: langchain.com/agents (optional framework)

✅ Final Checklist

Before deploying your agent, ensure:
☑ Max iterations set (prevent infinite loops)
☑ Timeout configured (prevent hanging)
☑ Error handling in place (graceful failures)
☑ Logging enabled (debug issues)
☑ Rate limiting active (protect APIs)
☑ Tool descriptions clear (guide LLM)
☑ Tests written (verify behavior)
☑ Monitoring setup (track performance)

Comments