Skip to main content

Integrating MCP with LLMs

Calculating read time…

Imagine you've built the perfect calculator tool using MCP. It can add, subtract, multiply, and divide with perfect precision.

But when you connect it to GPT-4 and ask "What's 5 + 3?", the AI responds with:
"5 + 3 equals 8"
without ever calling your calculator tool.

Frustrating, right? You built this amazing capability, but the AI just... ignored it.

This is the integration challenge. Building MCP tools is only half the battle. The real magic happens when you teach AI models to discover, understand, and correctly use your tools.

In this guide, you'll learn the complete integration process: how models find your tools, how to write prompts that trigger tool usage, how to prevent hallucinations, and how to build a working calculator that GPT actually uses!

🎯 The Integration Problem: Why AI Ignores Your Tools

Before diving into solutions, let's understand why AI models don't automatically use MCP tools.

The Three-Step Integration Dance

For an AI to successfully use your MCP tool, three things must happen:

Step 1: DISCOVERY AI learns: "These tools exist and here's what they do" Step 2: DECISION AI thinks: "Based on the user's request, I should use Tool X" Step 3: EXECUTION AI calls: Tool X with correct parameters AI receives: Result from tool AI responds: Incorporating the tool's output

Each step can fail independently, and understanding why is crucial.

Why Step 1 Fails: Discovery Problems

❌ Common Discovery Failures:

• Tools not properly advertised to the AI
• Poor tool descriptions (AI doesn't understand what the tool does)
• Missing or incorrect schema definitions
• Tool list too large (AI gets overwhelmed)

Why Step 2 Fails: Decision Problems

❌ Common Decision Failures:

• AI thinks it already knows the answer (doesn't need tools)
• User prompt doesn't clearly indicate tool usage is needed
• AI hallucinate an answer instead of using tools
• Tool description doesn't match user's intent

Why Step 3 Fails: Execution Problems

❌ Common Execution Failures:

• AI provides wrong parameter types or formats
• Missing required parameters
• Tool execution errors not handled gracefully
• AI doesn't know how to interpret tool results

The rest of this guide teaches you how to succeed at all three steps! 🎯

🔍 Step 1: Tool Discovery — How Models Find Your Tools

Let's start with the discovery process — how does an AI model learn about your MCP tools?

The Tool Discovery Flow

┌─────────────────────────────────────────┐ │ USER ASKS A QUESTION │ │ "What's 127 + 456?" │ └────────────────┬────────────────────────┘ │ ⬍ (1) Question sent to LLM │ ┌────────────────▼────────────────────────┐ │ LLM API CALL │ │ POST /v1/chat/completions │ │ │ │ { │ │ "model": "gpt-4", │ │ "messages": [...], │ │ "tools": [ │ │ { │ │ "type": "function", │ │ "function": { │ │ "name": "calculator", │ │ "description": "Add numbers", │ │ "parameters": {...} │ │ } │ │ } │ │ ] │ │ } │ └────────────────┬────────────────────────┘ │ ⬍ (2) LLM sees available tools │ ┌────────────────▼────────────────────────┐ │ LLM DECIDES TO USE TOOL │ │ "I should use calculator tool │ │ to add 127 + 456" │ └────────────────┬────────────────────────┘ │ ⬍ (3) Returns tool call │ ┌────────────────▼────────────────────────┐ │ TOOL CALL RESPONSE │ │ { │ │ "tool_calls": [{ │ │ "function": { │ │ "name": "calculator", │ │ "arguments": "{\"a\":127,\"b\":456,\"op\":\"add\"}" │ │ } │ │ }] │ │ } │ └─────────────────────────────────────────┘

Notice: Tools are passed to the LLM as part of the API request, NOT discovered by the model itself.

Converting MCP Tools to LLM Format

MCP tools need to be translated into the format that LLMs expect. Different LLM providers use slightly different formats.

MCP Tool Definition:

@mcp.tool()
def add_numbers(a: int, b: int) -> int:
    """Add two numbers together"""
    return a + b

OpenAI Format (Function Calling):

{
  "type": "function",
  "function": {
    "name": "add_numbers",
    "description": "Add two numbers together",
    "parameters": {
      "type": "object",
      "properties": {
        "a": {
          "type": "integer",
          "description": "First number"
        },
        "b": {
          "type": "integer", 
          "description": "Second number"
        }
      },
      "required": ["a", "b"]
    }
  }
}

Anthropic Format (Claude Tools):

{
  "name": "add_numbers",
  "description": "Add two numbers together",
  "input_schema": {
    "type": "object",
    "properties": {
      "a": {
        "type": "integer",
        "description": "First number"
      },
      "b": {
        "type": "integer",
        "description": "Second number"
      }
    },
    "required": ["a", "b"]
  }
}
Good News:

Libraries like LiteLLM and mcp-use handle this conversion automatically!
You don't need to manually convert between formats.

Writing Effective Tool Descriptions

The description is the most critical part of tool discovery. It tells the AI when and how to use your tool.

❌ Bad Tool Description:

"A calculator"

Problems:
• Too vague
• Doesn't specify what operations it supports
• AI might not know when to use it
Good Tool Description:

"Performs precise arithmetic operations (addition, subtraction, multiplication, division) on two numbers. Use this tool whenever you need to calculate exact mathematical results instead of estimating."

Why it works:
• Specifies exact operations
• Clear trigger: "whenever you need to calculate"
• Emphasizes precision vs estimation

Parameter Descriptions Matter Too

❌ Bad Parameter Descriptions:

a: "number"
b: "number"

AI might confuse which is which
✅ Good Parameter Descriptions:

a: "The first number in the operation"
b: "The second number in the operation"
operation: "The operation to perform: 'add', 'subtract', 'multiply', or 'divide'"

Clear, specific, unambiguous

🧠 Step 2: Getting AI to DECIDE to Use Tools

Even with perfect tool definitions, AI might not use them. You need to craft prompts that trigger tool usage.

The Hallucination Problem

Here's what often happens:

User: "What's 17 * 23?"

AI (without using calculator): 
"17 * 23 = 391"

[WRONG! Actual answer: 391]

Why did this happen? The AI hallucinated the answer instead of using the calculator tool.

💡 Why AI Hallucinates Instead of Using Tools:

1. Overconfidence: AI thinks it knows the answer
2. Pattern matching: Seen similar math in training data
3. No explicit instruction: User didn't demand precision
4. Tool cost: Using tools is "expensive" (extra tokens)

Tool-Friendly Prompting Techniques

Let's learn prompting patterns that encourage tool usage.

Technique 1: Explicit Tool Instructions

❌ Weak Prompt:

"What's 17 * 23?"

AI likely hallucinates the answer
✅ Strong Prompt:

"Use the calculator tool to compute 17 * 23"

Direct instruction forces tool usage

Technique 2: Demand Precision

✅ Precision-Focused Prompt:

"I need the EXACT result of 17 * 23. Do not estimate — use tools to calculate."

"EXACT" triggers tool usage

Technique 3: System Prompts for Tool Preference

Add a system message that encourages tool usage:

system_message = """
You are a helpful assistant with access to calculation tools.
When performing any mathematical computation:
1. ALWAYS use the calculator tool
2. NEVER estimate or calculate mentally
3. Show your work by explaining what you're calculating

This ensures 100% accuracy in all mathematical operations.
"""

Technique 4: Tool Choice Parameter

Most LLM APIs support a tool_choice parameter:

# Force tool usage
tool_choice = "required"  # Must call a tool

# Prefer specific tool
tool_choice = {"type": "function", "function": {"name": "calculator"}}

# Let AI decide (default)
tool_choice = "auto"
✅ Best Practice:

For critical operations (payments, calculations, database updates):
Use tool_choice = "required"

For general assistance:
Use tool_choice = "auto"

Multi-Step Reasoning with Tools

Sometimes AI needs to use tools multiple times:

User: "If I have $100 and earn 5% interest monthly for 3 months, how much will I have?"

Step 1: Calculate first month
Tool call: calculate(100, 1.05, "multiply") → 105

Step 2: Calculate second month  
Tool call: calculate(105, 1.05, "multiply") → 110.25

Step 3: Calculate third month
Tool call: calculate(110.25, 1.05, "multiply") → 115.76

Final answer: "You'll have $115.76"

This requires the AI to chain tool calls together.

💡 Enabling Multi-Step Tool Usage:

• Use models that support multiple tool calls per turn (GPT-4, Claude 3+)
• System prompt: "Break complex problems into steps, using tools for each calculation"
• Some APIs support parallel_tool_calls for simultaneous execution

🛡️ Preventing Hallucinations with Tools

Let's tackle the hallucination problem systematically.

The Three Types of Tool-Related Hallucinations

Type 1: Ignoring Tools (Most Common)

AI gives an answer without using available tools.

Solution:

  • Use tool_choice = "required"
  • Add "You MUST use tools" to system prompt
  • Verify tool calls in application logic

Type 2: Hallucinating Tool Results

AI pretends to call a tool but makes up the result.

AI: "I used the calculator tool and got 391"
[No tool was actually called!]

Solution:

  • Validate that tool_calls array is non-empty in response
  • Only trust results that come from actual tool execution
  • Log all tool calls for auditing

Type 3: Misusing Tools

AI calls tools with wrong parameters or wrong tool for the task.

Solution:

  • Strong parameter descriptions
  • Parameter validation in tool implementation
  • Return clear error messages

Validation Checklist

✅ Tool Usage Validation Pattern:

def validate_tool_usage(response):
    # Check 1: Was a tool called?
    if not response.get("tool_calls"):
        return False, "No tool was called"
    
    # Check 2: Are parameters valid?
    for call in response["tool_calls"]:
        args = json.loads(call["function"]["arguments"])
        if not validate_parameters(args):
            return False, "Invalid parameters"
    
    # Check 3: Did we get results?
    if not has_tool_results(response):
        return False, "Tool execution failed"
    
    return True, "Valid tool usage"

Complete Example: Calculator with GPT-4

Now let's build a complete working example! We'll create an MCP calculator server and connect it to GPT-4.

Step 1: Create the MCP Calculator Server

Create calculator_server.py:

"""
MCP Calculator Server
A simple calculator that GPT can use for accurate math
"""

from mcp.server.fastmcp import FastMCP

# Create MCP server
mcp = FastMCP("Calculator Server")

# ============================================
# CALCULATOR TOOL
# ============================================

@mcp.tool()
def add(a: float, b: float) -> dict:
    """
    Add two numbers with perfect precision.
    Use this tool whenever you need to add numbers accurately.
    Do not estimate - always use this tool for addition.
    
    Args:
        a: The first number to add
        b: The second number to add
    
    Returns:
        Dictionary with the sum and operation details
    """
    result = a + b
    return {
        "operation": "addition",
        "a": a,
        "b": b,
        "result": result,
        "formula": f"{a} + {b} = {result}"
    }

@mcp.tool()
def subtract(a: float, b: float) -> dict:
    """
    Subtract the second number from the first with perfect precision.
    Use this tool for accurate subtraction.
    
    Args:
        a: The number to subtract from
        b: The number to subtract
    
    Returns:
        Dictionary with the difference and operation details
    """
    result = a - b
    return {
        "operation": "subtraction",
        "a": a,
        "b": b,
        "result": result,
        "formula": f"{a} - {b} = {result}"
    }

@mcp.tool()
def multiply(a: float, b: float) -> dict:
    """
    Multiply two numbers with perfect precision.
    Use this tool for accurate multiplication.
    Do not estimate - always use this tool for multiplication.
    
    Args:
        a: The first number to multiply
        b: The second number to multiply
    
    Returns:
        Dictionary with the product and operation details
    """
    result = a * b
    return {
        "operation": "multiplication",
        "a": a,
        "b": b,
        "result": result,
        "formula": f"{a} × {b} = {result}"
    }

@mcp.tool()
def divide(a: float, b: float) -> dict:
    """
    Divide the first number by the second with perfect precision.
    Use this tool for accurate division.
    Handles division by zero safely.
    
    Args:
        a: The number to be divided (numerator)
        b: The number to divide by (denominator)
    
    Returns:
        Dictionary with the quotient and operation details
    """
    if b == 0:
        return {
            "operation": "division",
            "a": a,
            "b": b,
            "result": None,
            "error": "Cannot divide by zero",
            "formula": f"{a} ÷ {b} = undefined"
        }
    
    result = a / b
    return {
        "operation": "division",
        "a": a,
        "b": b,
        "result": result,
        "formula": f"{a} ÷ {b} = {result}"
    }

if __name__ == "__main__":
    print("Calculator MCP Server Starting...")
    print("Available tools: add, subtract, multiply, divide")
    mcp.run()

Step 2: Install Required Packages

pip install mcp litellm openai python-dotenv

Step 3: Create the Integration Script

Create use_calculator.py:

"""
Connect MCP Calculator to GPT-4
Demonstrates complete MCP-LLM integration
"""

import asyncio
import json
import os
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm
from dotenv import load_dotenv

# Load environment variables
load_dotenv()

async def main():
    print(" Starting MCP-GPT Integration Demo\n")
    
    # ============================================
    # STEP 1: Connect to MCP Server
    # ============================================
    
    print(" Connecting to MCP Calculator Server...")
    
    server_params = StdioServerParameters(
        command="python",
        args=["calculator_server.py"]
    )
    
    async with stdio_client(server_params) as (read, write):
        async with ClientSession(read, write) as session:
            
            # Initialize the MCP session
            await session.initialize()
            print(" Connected to MCP Server\n")
            
            # ============================================
            # STEP 2: Load MCP Tools
            # ============================================
            
            print("🔧 Loading MCP tools...")
            
            # Convert MCP tools to OpenAI format
            tools = await experimental_mcp_client.load_mcp_tools(
                session=session,
                format="openai"
            )
            
            print(f" Loaded {len(tools)} tools:")
            for tool in tools:
                print(f"   • {tool['function']['name']}")
            print()
            
            # ============================================
            # STEP 3: Create System Prompt
            # ============================================
            
            system_prompt = """
You are a helpful assistant with access to calculation tools.

IMPORTANT RULES:
1. For ANY mathematical calculation, you MUST use the calculator tools
2. NEVER calculate mentally or estimate
3. ALWAYS show which tool you're using and the result
4. Explain your work step by step

Available tools:
- add(a, b): Add two numbers
- subtract(a, b): Subtract b from a  
- multiply(a, b): Multiply two numbers
- divide(a, b): Divide a by b

Use these tools to ensure 100% accuracy.
            """.strip()
            
            # ============================================
            # STEP 4: Test Cases
            # ============================================
            
            test_cases = [
                {
                    "query": "What is 127 + 456?",
                    "expected_tool": "add",
                    "expected_result": 583
                },
                {
                    "query": "Calculate 17 * 23",
                    "expected_tool": "multiply", 
                    "expected_result": 391
                },
                {
                    "query": "If I have 1000 and spend 347, how much is left?",
                    "expected_tool": "subtract",
                    "expected_result": 653
                },
                {
                    "query": "Divide 144 by 12",
                    "expected_tool": "divide",
                    "expected_result": 12
                }
            ]
            
            # ============================================
            # STEP 5: Run Tests
            # ============================================
            
            for i, test in enumerate(test_cases, 1):
                print(f"{'='*60}")
                print(f"TEST {i}: {test['query']}")
                print(f"{'='*60}\n")
                
                messages = [
                    {"role": "system", "content": system_prompt},
                    {"role": "user", "content": test['query']}
                ]
                
                # Call GPT with tools
                response = await litellm.acompletion(
                    model="gpt-4o",
                    api_key=os.getenv("OPENAI_API_KEY"),
                    messages=messages,
                    tools=tools,
                    tool_choice="auto"  # Let AI decide when to use tools
                )
                
                # Extract response
                message = response.choices[0].message
                
                # Check if tool was called
                if message.tool_calls:
                    print("🔧 Tool Called!")
                    
                    for tool_call in message.tool_calls:
                        tool_name = tool_call.function.name
                        tool_args = json.loads(tool_call.function.arguments)
                        
                        print(f"   Tool: {tool_name}")
                        print(f"   Arguments: {tool_args}")
                        
                        # Execute the tool via MCP
                        result = await session.call_tool(
                            name=tool_name,
                            arguments=tool_args
                        )
                        
                        # Parse result
                        result_data = json.loads(result.content[0].text)
                        print(f"   Result: {result_data['result']}")
                        print(f"   Formula: {result_data['formula']}")
                        
                        # Verify correctness
                        if result_data['result'] == test['expected_result']:
                            print("    CORRECT!")
                        else:
                            print(f"    WRONG! Expected {test['expected_result']}")
                        
                        # Send result back to GPT for final response
                        messages.append({
                            "role": "assistant",
                            "content": None,
                            "tool_calls": message.tool_calls
                        })
                        
                        messages.append({
                            "role": "tool",
                            "tool_call_id": tool_call.id,
                            "content": json.dumps(result_data)
                        })
                        
                        # Get final response
                        final_response = await litellm.acompletion(
                            model="gpt-4o",
                            api_key=os.getenv("OPENAI_API_KEY"),
                            messages=messages,
                            tools=tools
                        )
                        
                        final_message = final_response.choices[0].message.content
                        print(f"\n💬 GPT's Final Answer:")
                        print(f"   {final_message}")
                        
                else:
                    print("  No tool called!")
                    print(f"   AI Response: {message.content}")
                    print("    AI should have used a tool!")
                
                print()
            
            print(f"{'='*60}")
            print(" Integration Demo Complete!")
            print(f"{'='*60}")

if __name__ == "__main__":
    asyncio.run(main())

Step 4: Create .env File

Create a .env file with your OpenAI API key:

OPENAI_API_KEY=sk-your-actual-api-key-here

Step 5: Run the Integration!

python use_calculator.py

You should see output like:

 Starting MCP-GPT Integration Demo

 Connecting to MCP Calculator Server...
 Connected to MCP Server

🔧 Loading MCP tools...
   Loaded 4 tools:
   • add
   • subtract
   • multiply
   • divide

============================================================
TEST 1: What is 127 + 456?
============================================================

🔧 Tool Called!
   Tool: add
   Arguments: {'a': 127, 'b': 456}
   Result: 583
   Formula: 127 + 456 = 583
   ✅ CORRECT!

💬 GPT's Final Answer:
   The sum of 127 and 456 is 583.

============================================================
TEST 2: Calculate 17 * 23
============================================================

🔧 Tool Called!
   Tool: multiply
   Arguments: {'a': 17, 'b': 23}
   Result: 391
   Formula: 17 × 23 = 391
   ✅ CORRECT!

💬 GPT's Final Answer:
   17 multiplied by 23 equals 391.
Congratulations!

You just successfully integrated MCP tools with GPT-4!
The AI discovered your tools, decided to use them, and executed them correctly.

🎯 Your Exercise: Add a Power Function

Now it's your turn to extend the calculator!

Challenge:

  1. Add a power tool to calculator_server.py
  2. It should calculate a^b (a raised to power b)
  3. Test it by asking GPT: "What's 2 to the power of 10?"

Solution Template:

@mcp.tool()
def power(base: float, exponent: float) -> dict:
    """
    Calculate base raised to the power of exponent.
    Use this tool for accurate exponentiation.
    
    Args:
        base: The base number
        exponent: The power to raise the base to
    
    Returns:
        Dictionary with the result and operation details
    """
    result = base ** exponent
    return {
        "operation": "exponentiation",
        "base": base,
        "exponent": exponent,
        "result": result,
        "formula": f"{base}^{exponent} = {result}"
    }

Test Your Implementation:

Add this test case to use_calculator.py:

{
    "query": "What's 2 to the power of 10?",
    "expected_tool": "power",
    "expected_result": 1024
}

Run again and verify GPT uses your new tool!

Advanced Integration Patterns

Pattern 1: Tool Chaining

Let AI use multiple tools in sequence:

User: "Calculate (5 + 3) * (10 - 2)"

Step 1: add(5, 3) → 8
Step 2: subtract(10, 2) → 8  
Step 3: multiply(8, 8) → 64

Answer: 64

Implementation tip:

system_prompt = """
For complex calculations:
1. Break them into steps
2. Use one tool per step
3. Show your work
4. Combine the results
"""

Pattern 2: Conditional Tool Usage

Use different tools based on input:

@mcp.tool()
def smart_calculator(expression: str) -> dict:
    """
    Evaluates mathematical expressions intelligently.
    Detects the operation and uses the appropriate calculation method.
    """
    if '+' in expression:
        # Parse and use add tool
        ...
    elif '*' in expression:
        # Parse and use multiply tool
        ...

Pattern 3: Tool Results as Context

Use tool results to inform next steps:

User: "I have $1000. If I invest at 5% yearly, how much after 10 years?"

Step 1: Get current amount → 1000
Step 2: Calculate growth → 1000 * (1.05 ^ 10)
Step 3: Use power tool → 1.629
Step 4: Use multiply tool → 1629

Answer: $1,629 after 10 years

🛠️ Debugging Integration Issues

Issue 1: Tools Not Appearing

⚠️ Symptom:
Loaded 0 tools

Solutions:
• Check MCP server is running
• Verify @mcp.tool() decorators are present
• Ensure session.initialize() was called
• Check server logs for errors

Issue 2: AI Not Using Tools

⚠️ Symptom:
AI answers without calling tools

Solutions:
• Use tool_choice="required"
• Add "You MUST use tools" to system prompt
• Make tool descriptions more specific
• Rephrase user query to be more explicit

Issue 3: Wrong Parameters

⚠️ Symptom:
Tool called but with incorrect arguments

Solutions:
• Improve parameter descriptions
• Add type hints (int, float, str)
• Validate parameters in tool function
• Return helpful error messages

Debugging Checklist

✅ Integration Debug Checklist:

# 1. Verify tools are loaded
print(f"Tools loaded: {len(tools)}")
for tool in tools:
    print(tool['function']['name'])

# 2. Log every LLM call
print("Sending to LLM:", messages)

# 3. Check response for tool calls
if response.choices[0].message.tool_calls:
    print("✅ Tool called!")
else:
    print("❌ No tool called")

# 4. Validate tool execution
try:
    result = await session.call_tool(...)
    print("Tool result:", result)
except Exception as e:
    print(f"Tool error: {e}")

# 5. Trace full conversation
print("Full conversation:", messages)

📊 Best Practices Summary

Tool Descriptions:

✓ Be specific about what the tool does
✓ Include when to use it ("Use this when...")
✓ Describe parameters clearly
✓ Mention limitations or edge cases
Prompting:

✓ Use system prompts to set tool usage expectations
✓ Demand precision for critical operations
✓ Use tool_choice parameter strategically
✓ Break complex tasks into steps
Hallucination Prevention:

✓ Validate tool calls actually happened
✓ Verify parameters are correct
✓ Handle errors gracefully
✓ Log everything for debugging
Error Handling:

✓ Return structured error messages
✓ Suggest corrections when parameters are wrong
✓ Implement retries for transient failures
✓ Never let tools crash silently

📚 Essential Resources

Libraries & Tools:
• LiteLLM: docs.litellm.ai
• MCP SDK: github.com/anthropics/mcp-python
• OpenAI Cookbook: cookbook.openai.com
• mcp-use framework: github.com/mcp-use/mcp-use

Documentation:
• OpenAI Function Calling: platform.openai.com/docs/guides/function-calling
• Claude Tool Use: docs.anthropic.com/claude/docs/tool-use
• MCP Spec: modelcontextprotocol.io

Community:
• MCP Discord Server
• OpenAI Developer Forum
• Stack Overflow #mcp tag

Comments