Imagine having an AI assistant that doesn't just answer questions, but actually does things for you.
It checks your calendar, sends emails, analyzes data, and even writes code — all by deciding which tools to use and when to use them.
This isn't science fiction. It's what happens when you combine Large Language Models (LLMs) with the Model Context Protocol (MCP) in an agent loop.
In this comprehensive guide, you'll learn:
- How to let LLMs autonomously decide which tools to call
- How to route LLM outputs to MCP tools
- How to feed tool results back to the LLM
- How to build complete agent loops that can handle complex tasks
By the end, you'll build a fully functional AI agent from scratch using only Python.
💡 What You'll Build
A complete research assistant that can search the web, analyze data, and write reports —
all by orchestrating multiple tools autonomously.
No frameworks, just pure Python. You'll understand every line.
🎯 Part 1: The Foundation — Understanding Agent Loops
Before we write code, let's understand what an agent loop actually is.
What is an Agent Loop?
An agent loop is a repeating cycle where an AI makes decisions, takes actions, observes results, and repeats.
Think of it like a human solving a problem:
- Think: "What should I do next?"
- Act: Take an action (use a tool)
- Observe: See what happened
- Reflect: Did that help? What next?
- Repeat: Keep going until the task is complete
This is called the ReAct pattern (Reasoning + Acting).
The Simple Analogy 🧭
Imagine you're cooking a new recipe but don't have all the ingredients.
Your thought process:
- "I need eggs" → Check the fridge (tool call)
- "No eggs found" → Go to the store (another tool call)
- "Bought eggs" → Return home (action complete)
- "Now I can cook" → Start cooking (final action)
Each step depends on the previous one. You don't plan everything upfront — you adapt based on what you discover.
That's exactly how an agent loop works.
The Agent Loop Architecture
Here's the complete flow:
┌─────────────────────────────────────────────┐
│ START: User Query │
└────────────────┬────────────────────────────┘
│
▼
┌─────────────────────────────────────────────┐
│ LLM: Analyze query and decide action │
│ - Can I answer directly? │
│ - Do I need a tool? │
│ - Which tool should I use? │
└────────────────┬────────────────────────────┘
│
┌────────┴────────┐
│ │
▼ ▼
┌─────────┐ ┌──────────────┐
│ Answer │ │ Tool Call │
│ Directly│ │ Request │
└────┬────┘ └──────┬───────┘
│ │
│ ▼
│ ┌─────────────────┐
│ │ MCP Server │
│ │ Execute Tool │
│ └────────┬────────┘
│ │
│ ▼
│ ┌─────────────────┐
│ │ Tool Result │
│ └────────┬────────┘
│ │
│ ▼
│ ┌─────────────────┐
│ │ Feed Result │
│ │ Back to LLM │
│ └────────┬────────┘
│ │
│ │
└──────────────────┴─────► LOOP CONTINUES
(until task complete)
The loop continues until the LLM decides it has enough information to answer the user.
🎯 Key Insight
The LLM doesn't execute tools — it generates requests to execute them.
Your code executes the tools and feeds results back to the LLM.
The LLM then decides if it needs more tools or if it can answer the user.
🔧 Part 2: Letting the LLM Decide Which Tool to Call
The first challenge: How does the LLM know which tools are available and when to use them?
Step 1: Define Your Tools
Let's create three simple tools for our agent:
# tools.py
import random
from datetime import datetime
def get_current_time():
"""Get the current date and time."""
return datetime.now().strftime("%Y-%m-%d %H:%M:%S")
def calculate(expression: str):
"""
Safely evaluate a mathematical expression.
Args:
expression: A math expression like "2 + 2" or "15 * 3"
Returns:
The calculated result
"""
try:
# Only allow safe math operations
allowed_chars = set('0123456789+-*/(). ')
if not all(c in allowed_chars for c in expression):
return "Error: Invalid characters in expression"
result = eval(expression)
return f"{expression} = {result}"
except Exception as e:
return f"Error: {str(e)}"
def get_random_fact():
"""Get a random interesting fact."""
facts = [
"Honey never spoils. Archaeologists have found 3000-year-old honey that's still edible.",
"A day on Venus is longer than its year.",
"Bananas are berries, but strawberries aren't.",
"There are more stars in the universe than grains of sand on Earth.",
]
return random.choice(facts)
Notice how each function has a clear docstring. This is crucial — the LLM reads these to understand what each tool does.
Step 2: Create Tool Descriptions for the LLM
We need to convert our Python functions into a format the LLM can understand.
# tool_registry.py
import inspect
import json
def get_tool_schema(func):
"""
Convert a Python function into a tool schema for the LLM.
"""
sig = inspect.signature(func)
params = {}
for name, param in sig.parameters.items():
param_type = "string" # Default type
if param.annotation != inspect.Parameter.empty:
if param.annotation == int:
param_type = "integer"
elif param.annotation == float:
param_type = "number"
elif param.annotation == bool:
param_type = "boolean"
params[name] = {
"type": param_type,
"description": f"Parameter: {name}"
}
return {
"name": func.__name__,
"description": func.__doc__ or f"Function {func.__name__}",
"parameters": {
"type": "object",
"properties": params,
"required": list(params.keys())
}
}
# Register all available tools
TOOLS = {
"get_current_time": get_current_time,
"calculate": calculate,
"get_random_fact": get_random_fact
}
# Generate schemas
TOOL_SCHEMAS = [get_tool_schema(func) for func in TOOLS.values()]
def get_tools_description():
"""Get a formatted description of all tools for the LLM."""
descriptions = []
for schema in TOOL_SCHEMAS:
params = schema["parameters"]["properties"]
param_desc = ", ".join([f"{k}: {v['type']}" for k, v in params.items()])
if param_desc:
desc = f"{schema['name']}({param_desc}) - {schema['description']}"
else:
desc = f"{schema['name']}() - {schema['description']}"
descriptions.append(desc)
return "\n".join(descriptions)
This code does something powerful: it automatically generates tool descriptions from your Python functions.
Step 3: Build the LLM Interface
Now we need to communicate with an LLM. We'll use OpenAI's API, but the same pattern works with any LLM.
# llm_client.py
import openai
import os
import json
openai.api_key = os.getenv("OPENAI_API_KEY")
def ask_llm(messages, tools=None):
"""
Send messages to the LLM and get a response.
Args:
messages: List of conversation messages
tools: Optional list of tool schemas
Returns:
The LLM's response
"""
params = {
"model": "gpt-4",
"messages": messages,
"temperature": 0.7,
}
if tools:
params["tools"] = [{"type": "function", "function": tool} for tool in tools]
params["tool_choice"] = "auto" # Let LLM decide
response = openai.chat.completions.create(**params)
return response.choices[0].message
def parse_tool_calls(message):
"""
Extract tool calls from the LLM's response.
Returns:
List of tool calls with name and arguments
"""
if not message.tool_calls:
return []
tool_calls = []
for call in message.tool_calls:
tool_calls.append({
"id": call.id,
"name": call.function.name,
"arguments": json.loads(call.function.arguments)
})
return tool_calls
Here's what this does:
- Sends your conversation to the LLM
- Provides tool schemas so the LLM knows what's available
- Lets the LLM decide whether to use a tool (
tool_choice: "auto") - Parses any tool calls from the response
✅ How LLM Tool Selection Works
The LLM reads your tool descriptions and compares them to the user's query.
If the user asks "What time is it?" the LLM matches this to get_current_time()
If the user asks "What's 234 * 56?" the LLM matches this to calculate(expression)
The LLM generates a JSON tool call request, not the actual execution.
🔄 Part 3: Routing LLM Output to MCP Tools
Now comes the exciting part: actually executing the tools the LLM requested.
Understanding the Routing Process
When the LLM decides to use a tool, it returns something like this:
{
"id": "call_abc123",
"name": "calculate",
"arguments": {
"expression": "234 * 56"
}
}
Your job is to:
- Read this JSON
- Find the corresponding tool function
- Call it with the provided arguments
- Capture the result
Building the Tool Executor
# tool_executor.py
from tool_registry import TOOLS
import json
def execute_tool(tool_name, arguments):
"""
Execute a tool by name with given arguments.
Args:
tool_name: Name of the tool to execute
arguments: Dictionary of arguments for the tool
Returns:
The tool's result as a string
"""
if tool_name not in TOOLS:
return f"Error: Tool '{tool_name}' not found"
try:
tool_function = TOOLS[tool_name]
# Call the function with unpacked arguments
result = tool_function(**arguments)
return str(result)
except TypeError as e:
return f"Error: Invalid arguments for {tool_name}: {str(e)}"
except Exception as e:
return f"Error executing {tool_name}: {str(e)}"
def execute_tool_calls(tool_calls):
"""
Execute multiple tool calls and return results.
Args:
tool_calls: List of tool call objects
Returns:
List of results with tool call IDs
"""
results = []
for call in tool_calls:
tool_id = call["id"]
tool_name = call["name"]
arguments = call["arguments"]
print(f"\n🔧 Executing tool: {tool_name}")
print(f" Arguments: {json.dumps(arguments, indent=2)}")
result = execute_tool(tool_name, arguments)
print(f" Result: {result}\n")
results.append({
"tool_call_id": tool_id,
"role": "tool",
"name": tool_name,
"content": result
})
return results
This is your router — it takes LLM output and executes the right tools.
Example Flow
Let's trace what happens step by step:
User asks: "What's 15 times 23?"
- LLM receives: The question + list of available tools
- LLM thinks: "This is a math question. I should use the calculate tool."
-
LLM returns:
{ "name": "calculate", "arguments": {"expression": "15 * 23"} } -
Your code executes:
calculate("15 * 23") -
Tool returns:
"15 * 23 = 345"
But we're not done! We need to feed this result back to the LLM.
🔁 Part 4: Feeding Tool Results Back to the LLM
After executing a tool, we need to give the result back to the LLM so it can continue reasoning.
The Feedback Loop
Here's how it works:
- User asks a question
- LLM requests a tool
- You execute the tool
- You append the result to the conversation
- You send the updated conversation back to the LLM
- LLM reads the result and decides what to do next
The key is maintaining a conversation history that includes:
- User messages
- LLM responses
- Tool calls
- Tool results
Building the Conversation Manager
# conversation.py
class Conversation:
"""Manages the conversation history between user, LLM, and tools."""
def __init__(self, system_prompt):
self.messages = [
{"role": "system", "content": system_prompt}
]
def add_user_message(self, content):
"""Add a user message to the conversation."""
self.messages.append({
"role": "user",
"content": content
})
def add_assistant_message(self, message):
"""Add an LLM response to the conversation."""
msg = {
"role": "assistant",
"content": message.content or ""
}
# Include tool calls if present
if message.tool_calls:
msg["tool_calls"] = [
{
"id": call.id,
"type": "function",
"function": {
"name": call.function.name,
"arguments": call.function.arguments
}
}
for call in message.tool_calls
]
self.messages.append(msg)
def add_tool_results(self, results):
"""Add tool execution results to the conversation."""
for result in results:
self.messages.append(result)
def get_messages(self):
"""Get the complete conversation history."""
return self.messages
def print_conversation(self):
"""Print the conversation in a readable format."""
for msg in self.messages:
role = msg["role"].upper()
content = msg.get("content", "")
if role == "SYSTEM":
print(f"\n{'='*50}")
print(f"SYSTEM: {content}")
print(f"{'='*50}\n")
elif role == "USER":
print(f"\n👤 USER: {content}\n")
elif role == "ASSISTANT":
if content:
print(f"🤖 ASSISTANT: {content}\n")
if "tool_calls" in msg:
print(" [Requesting tool calls...]")
elif role == "TOOL":
print(f"🔧 TOOL ({msg['name']}): {content}\n")
This class keeps track of everything that happens in the conversation.
💡 Why Conversation History Matters
The LLM is stateless — it doesn't remember previous interactions.
By sending the complete conversation history each time, the LLM has context.
This allows it to make informed decisions based on what tools were called and what results came back.
♾️ Part 5: Building Complete Agent Loops
Now we combine everything into a full agent loop that can handle complex multi-step tasks.
The Agent Loop Implementation
# agent.py
from llm_client import ask_llm, parse_tool_calls
from tool_executor import execute_tool_calls
from tool_registry import TOOL_SCHEMAS, get_tools_description
from conversation import Conversation
class Agent:
"""
An autonomous agent that uses LLM + MCP tools to solve tasks.
"""
def __init__(self, max_iterations=10):
"""
Initialize the agent.
Args:
max_iterations: Maximum number of loop iterations to prevent infinite loops
"""
self.max_iterations = max_iterations
# Create system prompt with tool descriptions
tools_desc = get_tools_description()
self.system_prompt = f"""You are a helpful AI assistant with access to tools.
Available tools:
{tools_desc}
When you need to use a tool, request it and wait for the result.
When you have enough information, provide a final answer to the user.
Be concise and helpful."""
def run(self, user_query):
"""
Run the agent loop for a user query.
Args:
user_query: The user's question or task
Returns:
The final answer
"""
print(f"\n{'='*60}")
print(f"🚀 Starting Agent Loop")
print(f"{'='*60}\n")
print(f"👤 User Query: {user_query}\n")
# Initialize conversation
conversation = Conversation(self.system_prompt)
conversation.add_user_message(user_query)
iteration = 0
# Main agent loop
while iteration < self.max_iterations:
iteration += 1
print(f"\n--- Iteration {iteration} ---\n")
# Get LLM response
llm_response = ask_llm(
conversation.get_messages(),
tools=TOOL_SCHEMAS
)
# Add LLM response to conversation
conversation.add_assistant_message(llm_response)
# Check if LLM wants to use tools
tool_calls = parse_tool_calls(llm_response)
if not tool_calls:
# No tool calls - LLM has final answer
final_answer = llm_response.content
print(f"\n{'='*60}")
print(f"✅ Final Answer:")
print(f"{'='*60}\n")
print(f"{final_answer}\n")
return final_answer
# Execute tools
print(f"🔧 Executing {len(tool_calls)} tool(s)...\n")
tool_results = execute_tool_calls(tool_calls)
# Feed results back to conversation
conversation.add_tool_results(tool_results)
# Loop continues...
# Max iterations reached
print("\n⚠️ Max iterations reached. Stopping loop.\n")
return "I couldn't complete the task within the iteration limit."
# Example usage
if __name__ == "__main__":
agent = Agent(max_iterations=5)
# Example 1: Simple calculation
agent.run("What's 456 multiplied by 789?")
# Example 2: Multiple tools
agent.run("What time is it, and can you also tell me a random fact?")
This is a complete, working agent! Let's break down what happens:
- User provides a query
- Agent sends query + tool schemas to LLM
- LLM decides if it needs tools
- If yes → Agent executes tools and feeds results back (loop continues)
- If no → LLM provides final answer (loop ends)
Example Agent Execution
Let's trace a real example: "What's 234 * 56 and what time is it?"
Iteration 1:
LLM thinks: "User wants calculation AND current time. I need two tools." LLM requests: 1. calculate(expression="234 * 56") 2. get_current_time() Agent executes both tools: - calculate returns: "234 * 56 = 13104" - get_current_time returns: "2026-02-04 14:30:25" Agent feeds results back to LLM...
Iteration 2:
LLM reads tool results. LLM thinks: "I now have both pieces of information. I can answer." LLM responds: "The calculation 234 * 56 equals 13,104. The current time is 2:30 PM on February 4th, 2026." No tool calls → Loop ends with final answer.
The agent handled a multi-step task completely autonomously!
✅ Agent Loop Best Practices
1. Set max iterations to prevent infinite loops
2. Log everything for debugging
3. Handle errors gracefully in tool execution
4. Give clear tool descriptions to guide LLM
5. Include system prompts that explain the agent's role
🎯 Part 6: Advanced Agent Patterns
Now that you understand the basics, let's explore advanced patterns.
Pattern 1: Conditional Tool Calling
Sometimes you want the LLM to always use a specific tool first.
# Force tool usage on first call
llm_response = ask_llm(
messages,
tools=TOOL_SCHEMAS,
tool_choice={"type": "function", "function": {"name": "search_database"}}
)
This forces the LLM to call search_database first, then decide what to do with the results.
Pattern 2: Tool Chaining
Build tools that depend on other tools' outputs.
def search_products(query: str):
"""Search for products matching a query."""
# Returns list of product IDs
return ["prod_123", "prod_456", "prod_789"]
def get_product_details(product_id: str):
"""Get detailed information about a specific product."""
# Returns product details
return {
"id": product_id,
"name": "Example Product",
"price": 99.99
}
The LLM can chain these:
- Call
search_products("laptop")→ Get IDs - Call
get_product_details("prod_123")→ Get details - Analyze details and answer user
Pattern 3: Parallel Tool Execution
When tools don't depend on each other, execute them in parallel for speed.
import asyncio
async def execute_tool_async(tool_name, arguments):
"""Execute a single tool asynchronously."""
# Your tool execution logic
return execute_tool(tool_name, arguments)
async def execute_tools_parallel(tool_calls):
"""Execute multiple tools in parallel."""
tasks = [
execute_tool_async(call["name"], call["arguments"])
for call in tool_calls
]
results = await asyncio.gather(*tasks)
return results
If the LLM requests weather for 3 different cities, run all 3 API calls simultaneously.
Pattern 4: Memory and State Management
For long-running agents, add memory storage.
class StatefulAgent(Agent):
"""Agent with persistent memory."""
def __init__(self):
super().__init__()
self.memory = {} # Store information across runs
def store_memory(self, key, value):
"""Store information for future reference."""
self.memory[key] = value
def recall_memory(self, key):
"""Recall previously stored information."""
return self.memory.get(key)
Now tools can save and retrieve information across multiple user queries.
🏗️ Part 7: Building a Real-World Research Assistant
Let's build a complete research assistant that can search, analyze, and report.
Define Research Tools
# research_tools.py
import requests
from bs4 import BeautifulSoup
def web_search(query: str, num_results: int = 3):
"""
Search the web and return top results.
Args:
query: Search query
num_results: Number of results to return
Returns:
List of search results with titles and URLs
"""
# In production, use a real search API like Google Custom Search
# This is a simplified example
return [
{
"title": f"Result {i+1} for '{query}'",
"url": f"https://example.com/result{i+1}",
"snippet": f"Information about {query}..."
}
for i in range(num_results)
]
def fetch_webpage(url: str):
"""
Fetch and extract text content from a webpage.
Args:
url: The URL to fetch
Returns:
Extracted text content
"""
try:
response = requests.get(url, timeout=10)
soup = BeautifulSoup(response.content, 'html.parser')
# Extract text from paragraphs
paragraphs = soup.find_all('p')
text = ' '.join([p.get_text() for p in paragraphs[:10]]) # First 10 paragraphs
return text[:1000] # Limit to 1000 characters
except Exception as e:
return f"Error fetching webpage: {str(e)}"
def analyze_data(data: str, analysis_type: str):
"""
Analyze text data and extract insights.
Args:
data: The text data to analyze
analysis_type: Type of analysis ('summary', 'keywords', 'sentiment')
Returns:
Analysis results
"""
if analysis_type == "summary":
# Simplified summary - in production use NLP models
words = data.split()
return f"Data contains {len(words)} words. " + " ".join(words[:50]) + "..."
elif analysis_type == "keywords":
# Extract most common words
words = data.lower().split()
word_freq = {}
for word in words:
if len(word) > 4: # Only words longer than 4 chars
word_freq[word] = word_freq.get(word, 0) + 1
top_words = sorted(word_freq.items(), key=lambda x: x[1], reverse=True)[:5]
return f"Top keywords: {', '.join([w[0] for w in top_words])}"
elif analysis_type == "sentiment":
# Simplified sentiment - in production use sentiment analysis models
positive_words = ["good", "great", "excellent", "amazing", "wonderful"]
negative_words = ["bad", "terrible", "awful", "poor", "horrible"]
text_lower = data.lower()
pos_count = sum(text_lower.count(word) for word in positive_words)
neg_count = sum(text_lower.count(word) for word in negative_words)
if pos_count > neg_count:
return f"Sentiment: Positive (pos: {pos_count}, neg: {neg_count})"
elif neg_count > pos_count:
return f"Sentiment: Negative (pos: {pos_count}, neg: {neg_count})"
else:
return f"Sentiment: Neutral (pos: {pos_count}, neg: {neg_count})"
return "Unknown analysis type"
def save_report(filename: str, content: str):
"""
Save research findings to a file.
Args:
filename: Name of the file
content: Content to save
Returns:
Confirmation message
"""
try:
with open(filename, 'w') as f:
f.write(content)
return f"Report saved to {filename}"
except Exception as e:
return f"Error saving report: {str(e)}"
Create the Research Agent
# research_agent.py
from agent import Agent
from tool_registry import TOOLS
from research_tools import web_search, fetch_webpage, analyze_data, save_report
# Register research tools
TOOLS.update({
"web_search": web_search,
"fetch_webpage": fetch_webpage,
"analyze_data": analyze_data,
"save_report": save_report
})
class ResearchAgent(Agent):
"""
Specialized agent for research tasks.
"""
def __init__(self):
super().__init__(max_iterations=15) # More iterations for complex research
# Enhanced system prompt for research
tools_desc = self.get_tools_description()
self.system_prompt = f"""You are an expert research assistant.
Your capabilities:
1. Search the web for information
2. Fetch and read webpage content
3. Analyze data for insights
4. Save research reports
Available tools:
{tools_desc}
Research Process:
1. Search for relevant information
2. Fetch detailed content from promising sources
3. Analyze the gathered data
4. Synthesize findings into a coherent report
5. Save the report if requested
Be thorough, accurate, and cite your sources."""
# Example usage
if __name__ == "__main__":
agent = ResearchAgent()
result = agent.run(
"Research the benefits of Python for data science and save a summary report to python_research.txt"
)
Example Research Session
Let's trace what happens when you ask: "Research AI safety and summarize the top 3 concerns"
Iteration 1: LLM calls web_search("AI safety concerns")
Iteration 2: LLM calls fetch_webpage(url1), fetch_webpage(url2), fetch_webpage(url3)
Iteration 3: LLM calls analyze_data(combined_text, "summary")
Iteration 4: LLM synthesizes all information and provides final answer with top 3 concerns
The agent orchestrated multiple tools across multiple steps completely autonomously!
🐛 Part 8: Error Handling and Recovery
Real-world agents need robust error handling.
Common Failure Modes
- Tool execution fails (API down, invalid input)
- LLM hallucinations (requests non-existent tools)
- Infinite loops (LLM keeps calling same tool)
- Timeout errors (slow APIs, network issues)
Enhanced Error Handling
# robust_agent.py
import time
from collections import Counter
class RobustAgent(Agent):
"""Agent with comprehensive error handling."""
def __init__(self, max_iterations=10, timeout_seconds=60):
super().__init__(max_iterations)
self.timeout_seconds = timeout_seconds
self.start_time = None
self.tool_call_history = []
def check_timeout(self):
"""Check if agent has exceeded time limit."""
elapsed = time.time() - self.start_time
if elapsed > self.timeout_seconds:
raise TimeoutError(f"Agent exceeded {self.timeout_seconds}s time limit")
def detect_loop(self):
"""Detect if agent is stuck in a loop."""
if len(self.tool_call_history) < 3:
return False
# Check last 3 calls
recent_calls = self.tool_call_history[-3:]
counter = Counter(recent_calls)
# If same tool called 3 times in a row, we're in a loop
if counter.most_common(1)[0][1] >= 3:
return True
return False
def execute_tool_safely(self, tool_name, arguments):
"""Execute tool with error handling and retry logic."""
max_retries = 3
retry_delay = 1 # seconds
for attempt in range(max_retries):
try:
result = execute_tool(tool_name, arguments)
# Track successful call
self.tool_call_history.append(tool_name)
return result
except Exception as e:
if attempt < max_retries - 1:
print(f"⚠️ Tool failed (attempt {attempt + 1}/{max_retries}): {str(e)}")
print(f" Retrying in {retry_delay}s...")
time.sleep(retry_delay)
retry_delay *= 2 # Exponential backoff
else:
error_msg = f"Tool {tool_name} failed after {max_retries} attempts: {str(e)}"
print(f"❌ {error_msg}")
return error_msg
def run(self, user_query):
"""Run agent with comprehensive error handling."""
self.start_time = time.time()
self.tool_call_history = []
try:
iteration = 0
conversation = Conversation(self.system_prompt)
conversation.add_user_message(user_query)
while iteration < self.max_iterations:
iteration += 1
# Check timeout
self.check_timeout()
# Check for loops
if self.detect_loop():
print("\n⚠️ Loop detected! Stopping to prevent infinite recursion.\n")
return "I seem to be stuck in a loop. Let me try a different approach."
# Get LLM response with error handling
try:
llm_response = ask_llm(conversation.get_messages(), tools=TOOL_SCHEMAS)
except Exception as e:
print(f"\n❌ LLM error: {str(e)}\n")
return f"I encountered an error while processing: {str(e)}"
conversation.add_assistant_message(llm_response)
tool_calls = parse_tool_calls(llm_response)
if not tool_calls:
return llm_response.content
# Execute tools safely
tool_results = []
for call in tool_calls:
result = self.execute_tool_safely(call["name"], call["arguments"])
tool_results.append({
"tool_call_id": call["id"],
"role": "tool",
"name": call["name"],
"content": result
})
conversation.add_tool_results(tool_results)
return "Max iterations reached without completing the task."
except TimeoutError as e:
return f"Operation timed out: {str(e)}"
except Exception as e:
return f"Unexpected error: {str(e)}"
This enhanced agent handles:
- Timeouts (prevents hanging)
- Infinite loops (detects repeated tool calls)
- Tool failures (retries with exponential backoff)
- LLM errors (graceful degradation)
❌ Common Pitfalls to Avoid
1. No max iterations: Agent runs forever
2. No error handling: One failed tool crashes everything
3. Vague tool descriptions: LLM doesn't know when to use them
4. Not logging: Impossible to debug what went wrong
5. Synchronous execution: Slow for parallel operations
📊 Part 9: Monitoring and Debugging
Production agents need monitoring to understand what's happening.
Adding Comprehensive Logging
# logger.py
import json
from datetime import datetime
class AgentLogger:
"""Logger for tracking agent execution."""
def __init__(self, log_file="agent_log.json"):
self.log_file = log_file
self.current_run = {
"start_time": None,
"end_time": None,
"user_query": None,
"iterations": [],
"final_answer": None,
"errors": []
}
def start_run(self, user_query):
"""Start logging a new agent run."""
self.current_run = {
"start_time": datetime.now().isoformat(),
"user_query": user_query,
"iterations": [],
"errors": []
}
def log_iteration(self, iteration_num, tool_calls, tool_results):
"""Log an iteration."""
self.current_run["iterations"].append({
"iteration": iteration_num,
"timestamp": datetime.now().isoformat(),
"tool_calls": tool_calls,
"results": tool_results
})
def log_error(self, error):
"""Log an error."""
self.current_run["errors"].append({
"timestamp": datetime.now().isoformat(),
"error": str(error)
})
def end_run(self, final_answer):
"""Complete the run log."""
self.current_run["end_time"] = datetime.now().isoformat()
self.current_run["final_answer"] = final_answer
# Save to file
with open(self.log_file, 'a') as f:
f.write(json.dumps(self.current_run, indent=2) + "\n")
def get_stats(self):
"""Get statistics about the run."""
return {
"total_iterations": len(self.current_run["iterations"]),
"total_tool_calls": sum(
len(it["tool_calls"]) for it in self.current_run["iterations"]
),
"errors": len(self.current_run["errors"]),
"duration": self._calculate_duration()
}
def _calculate_duration(self):
"""Calculate run duration."""
if not self.current_run["start_time"] or not self.current_run["end_time"]:
return None
start = datetime.fromisoformat(self.current_run["start_time"])
end = datetime.fromisoformat(self.current_run["end_time"])
return (end - start).total_seconds()
Visualization Dashboard
Create a simple dashboard to monitor agent performance:
# dashboard.py
import json
from collections import Counter
def analyze_logs(log_file="agent_log.json"):
"""Analyze agent logs and generate insights."""
with open(log_file, 'r') as f:
content = f.read()
runs = [json.loads(line) for line in content.strip().split('\n') if line]
total_runs = len(runs)
successful_runs = sum(1 for r in runs if r.get("final_answer") and not r.get("errors"))
all_tool_calls = []
for run in runs:
for iteration in run.get("iterations", []):
for call in iteration.get("tool_calls", []):
all_tool_calls.append(call["name"])
tool_usage = Counter(all_tool_calls)
avg_iterations = sum(len(r.get("iterations", [])) for r in runs) / total_runs if total_runs > 0 else 0
print("\n" + "="*60)
print("AGENT PERFORMANCE DASHBOARD")
print("="*60 + "\n")
print(f"Total Runs: {total_runs}")
print(f"Successful: {successful_runs} ({successful_runs/total_runs*100:.1f}%)")
print(f"Failed: {total_runs - successful_runs}")
print(f"\nAverage Iterations per Run: {avg_iterations:.1f}")
print(f"\nMost Used Tools:")
for tool, count in tool_usage.most_common(5):
print(f" {tool}: {count} calls")
print("\n" + "="*60 + "\n")
# Run analysis
if __name__ == "__main__":
analyze_logs()
🎓 Part 10: Best Practices and Optimization
Let's wrap up with production-ready best practices.
1. Optimize Token Usage
LLM API costs are based on tokens. Reduce costs by:
- Truncating long tool results (first 500 chars)
- Summarizing web content before feeding to LLM
- Using cheaper models for simple tool selection
- Caching tool results that don't change
2. Implement Caching
import hashlib
import json
class ToolCache:
"""Cache tool results to avoid redundant calls."""
def __init__(self):
self.cache = {}
def _get_key(self, tool_name, arguments):
"""Generate cache key from tool and args."""
key_str = f"{tool_name}:{json.dumps(arguments, sort_keys=True)}"
return hashlib.md5(key_str.encode()).hexdigest()
def get(self, tool_name, arguments):
"""Get cached result if available."""
key = self._get_key(tool_name, arguments)
return self.cache.get(key)
def set(self, tool_name, arguments, result):
"""Cache a tool result."""
key = self._get_key(tool_name, arguments)
self.cache[key] = result
# Use in agent
cache = ToolCache()
def execute_tool_with_cache(tool_name, arguments):
"""Execute tool with caching."""
cached = cache.get(tool_name, arguments)
if cached:
print(f"✅ Using cached result for {tool_name}")
return cached
result = execute_tool(tool_name, arguments)
cache.set(tool_name, arguments, result)
return result
3. Rate Limiting
import time
from collections import deque
class RateLimiter:
"""Rate limiter for API calls."""
def __init__(self, max_calls, time_window):
self.max_calls = max_calls
self.time_window = time_window # seconds
self.calls = deque()
def acquire(self):
"""Wait until rate limit allows next call."""
now = time.time()
# Remove old calls outside time window
while self.calls and self.calls[0] < now - self.time_window:
self.calls.popleft()
# Wait if at limit
if len(self.calls) >= self.max_calls:
sleep_time = self.calls[0] + self.time_window - now
if sleep_time > 0:
time.sleep(sleep_time)
self.acquire() # Retry
self.calls.append(now)
# 10 calls per minute
rate_limiter = RateLimiter(max_calls=10, time_window=60)
def rate_limited_llm_call(messages, tools):
"""Make LLM call with rate limiting."""
rate_limiter.acquire()
return ask_llm(messages, tools)
4. Testing Your Agent
# test_agent.py
import pytest
from agent import Agent
def test_simple_calculation():
"""Test agent can handle basic math."""
agent = Agent()
result = agent.run("What's 5 + 7?")
assert "12" in result
def test_multiple_tools():
"""Test agent can use multiple tools."""
agent = Agent()
result = agent.run("What time is it and what's 10 * 10?")
assert "100" in result # Calculation result
# Time check would need mocking
def test_error_handling():
"""Test agent handles tool errors gracefully."""
agent = Agent()
result = agent.run("Calculate 'invalid expression'")
# Should return error message, not crash
assert result is not None
def test_max_iterations():
"""Test agent stops at max iterations."""
agent = Agent(max_iterations=2)
# Complex query that would need many iterations
result = agent.run("Solve world hunger")
# Should stop after 2 iterations
assert "iteration limit" in result.lower() or result is not None
if __name__ == "__main__":
pytest.main([__file__])
🚀 Putting It All Together: Complete Production Agent
Here's a complete, production-ready agent with all best practices:
# production_agent.py
import time
from typing import Optional
from llm_client import ask_llm, parse_tool_calls
from tool_executor import execute_tool
from conversation import Conversation
from logger import AgentLogger
from tool_cache import ToolCache
from rate_limiter import RateLimiter
class ProductionAgent:
"""
Production-ready agent with full error handling, logging, caching, and monitoring.
"""
def __init__(
self,
tools,
tool_schemas,
max_iterations=10,
timeout_seconds=120,
enable_cache=True,
rate_limit_calls=60 # calls per minute
):
self.tools = tools
self.tool_schemas = tool_schemas
self.max_iterations = max_iterations
self.timeout_seconds = timeout_seconds
# Initialize components
self.logger = AgentLogger()
self.cache = ToolCache() if enable_cache else None
self.rate_limiter = RateLimiter(max_calls=rate_limit_calls, time_window=60)
# State tracking
self.tool_call_history = []
self.start_time = None
def run(self, user_query: str) -> str:
"""
Execute the agent loop with full production features.
"""
self.start_time = time.time()
self.logger.start_run(user_query)
print(f"\n{'='*60}")
print(f"🚀 Production Agent Starting")
print(f"{'='*60}\n")
print(f"Query: {user_query}\n")
try:
conversation = Conversation(self._get_system_prompt())
conversation.add_user_message(user_query)
iteration = 0
while iteration < self.max_iterations:
iteration += 1
print(f"\n--- Iteration {iteration} ---")
# Timeout check
if time.time() - self.start_time > self.timeout_seconds:
raise TimeoutError("Agent exceeded time limit")
# Loop detection
if self._detect_loop():
raise RuntimeError("Infinite loop detected")
# Get LLM response with rate limiting
self.rate_limiter.acquire()
llm_response = ask_llm(
conversation.get_messages(),
tools=self.tool_schemas
)
conversation.add_assistant_message(llm_response)
tool_calls = parse_tool_calls(llm_response)
# No tools needed - we have final answer
if not tool_calls:
final_answer = llm_response.content
self.logger.end_run(final_answer)
self._print_stats()
return final_answer
# Execute tools with caching and error handling
tool_results = self._execute_tools(tool_calls)
conversation.add_tool_results(tool_results)
# Log iteration
self.logger.log_iteration(iteration, tool_calls, tool_results)
# Max iterations reached
result = "Task incomplete: maximum iterations reached"
self.logger.end_run(result)
self._print_stats()
return result
except Exception as e:
self.logger.log_error(e)
error_msg = f"Error: {str(e)}"
self.logger.end_run(error_msg)
print(f"\n❌ {error_msg}\n")
return error_msg
def _execute_tools(self, tool_calls):
"""Execute tools with caching and retry logic."""
results = []
for call in tool_calls:
tool_name = call["name"]
arguments = call["arguments"]
# Try cache first
if self.cache:
cached = self.cache.get(tool_name, arguments)
if cached:
print(f"✅ Cache hit: {tool_name}")
results.append({
"tool_call_id": call["id"],
"role": "tool",
"name": tool_name,
"content": cached
})
continue
# Execute with retry
result = self._execute_with_retry(tool_name, arguments)
# Cache result
if self.cache:
self.cache.set(tool_name, arguments, result)
# Track for loop detection
self.tool_call_history.append(tool_name)
results.append({
"tool_call_id": call["id"],
"role": "tool",
"name": tool_name,
"content": result
})
return results
def _execute_with_retry(self, tool_name, arguments, max_retries=3):
"""Execute tool with exponential backoff retry."""
retry_delay = 1
for attempt in range(max_retries):
try:
return execute_tool(tool_name, arguments)
except Exception as e:
if attempt < max_retries - 1:
print(f"⚠️ Retry {attempt + 1}/{max_retries} for {tool_name}")
time.sleep(retry_delay)
retry_delay *= 2
else:
return f"Error after {max_retries} attempts: {str(e)}"
def _detect_loop(self):
"""Detect if agent is stuck in a loop."""
if len(self.tool_call_history) < 4:
return False
recent = self.tool_call_history[-4:]
return len(set(recent)) == 1 # Same tool 4 times
def _get_system_prompt(self):
"""Generate system prompt with tool descriptions."""
tools_desc = "\n".join([
f"- {s['name']}: {s['description']}"
for s in self.tool_schemas
])
return f"""You are a helpful AI assistant with access to tools.
Available tools:
{tools_desc}
Use tools when needed to accomplish tasks. When you have sufficient information, provide a final answer."""
def _print_stats(self):
"""Print execution statistics."""
stats = self.logger.get_stats()
print(f"\n{'='*60}")
print(f"📊 Execution Stats")
print(f"{'='*60}")
print(f"Iterations: {stats['total_iterations']}")
print(f"Tool calls: {stats['total_tool_calls']}")
print(f"Errors: {stats['errors']}")
print(f"Duration: {stats['duration']:.2f}s")
print(f"{'='*60}\n")
This production agent includes:
- ✅ Comprehensive error handling
- ✅ Tool result caching
- ✅ Rate limiting
- ✅ Timeout protection
- ✅ Loop detection
- ✅ Retry logic
- ✅ Complete logging
- ✅ Performance monitoring
Let's recap your complete journey:
Part 1: Understanding Agent Loops
- Agent loops follow the ReAct pattern (Reason → Act → Observe)
- LLMs decide actions, your code executes them
- Loop continues until task completion
Part 2: LLM Tool Selection
- Define tools with clear docstrings
- Convert functions to schemas the LLM can read
- Let LLM decide which tool to use based on user query
Part 3: Routing to MCP Tools
- Parse LLM's tool call requests (JSON format)
- Execute the requested tool with provided arguments
- Capture results for feedback
Part 4: Feedback Loop
- Maintain complete conversation history
- Append tool results to conversation
- Send updated history back to LLM
- LLM uses results to continue reasoning
Part 5: Production Features
- Error handling and retries
- Caching for efficiency
- Rate limiting for API protection
- Logging and monitoring
- Loop and timeout detection
📖 Additional Resources
- MCP Protocol Specification: modelcontextprotocol.io
- OpenAI Function Calling: platform.openai.com/docs/guides/function-calling
- Anthropic Claude Tool Use: docs.anthropic.com/claude/docs/tool-use
- ReAct Paper: Reason and Act pattern research
- LangChain Agents: langchain.com/agents (optional framework)
✅ Final Checklist
Before deploying your agent, ensure:
☑ Max iterations set (prevent infinite loops)
☑ Timeout configured (prevent hanging)
☑ Error handling in place (graceful failures)
☑ Logging enabled (debug issues)
☑ Rate limiting active (protect APIs)
☑ Tool descriptions clear (guide LLM)
☑ Tests written (verify behavior)
☑ Monitoring setup (track performance)
Comments
Post a Comment