Imagine you have a brilliant assistant who can do anything — search the web, check the weather, send emails, or even book flights.
But here's the catch: this assistant can't do any of these things directly. They need to call the right "tool" for each task.
The big question is: How does the assistant know which tool to use?
This is exactly what happens inside Large Language Models (LLMs) when they work with the Model Context Protocol (MCP). In this tutorial, we'll go from absolute zero to hero-level understanding of how LLMs make these decisions.
💡 Important Note
This article assumes zero prior knowledge. We'll build your understanding step by step with real examples.
🎬 Setting the Stage: What is MCP?
Before diving into tool selection, let's understand what MCP actually is.
Think of MCP as a universal translator between AI models and the outside world.
Without MCP, every time you wanted your AI to use a new tool (like a calendar app, database, or weather API), you'd need to write custom code to connect them. This creates a nightmare of custom integrations.
MCP solves this by providing a standardized protocol — like USB-C for AI applications.
The Simple Analogy 🔌
Remember when every phone had a different charger? Frustrating, right?
Then USB-C came along. One standard connector for everything.
That's MCP for AI. Instead of building custom connections for each tool, MCP provides one standard way to connect.
🎯 Key Takeaway
MCP is a protocol that lets LLMs discover and use external tools in a standardized way. But the LLM still needs to decide WHICH tool to use. That's what we're learning today.
🧠 Understanding the Core Concept: LLMs Don't Actually "Call" Anything
Here's a mind-bending truth that many people misunderstand:
LLMs don't actually execute functions or call tools themselves.
Wait, what? Let me explain.
What LLMs Actually Do
LLMs are fundamentally text prediction machines. They look at text and predict what comes next.
That's it. Nothing more, nothing less.
When an LLM "uses a tool," here's what's really happening:
- The LLM generates text that describes which tool to use and what parameters to pass
- Your application code reads that text and actually executes the function
- The result gets sent back to the LLM as more text
- The LLM generates a final response based on the result
Think of the LLM as a smart decision-maker that writes instructions, while your application is the executor that carries them out.
❌ Common Misconception
"The LLM calls the API and gets the data directly."
Reality: The LLM generates a structured request (JSON format), your code executes it,
and returns the result back to the LLM.
🔍 The Tool Selection Process: Step by Step
Now let's dive into the fascinating process of how an LLM chooses which tool to call.
We'll break this down into clear stages with real examples.
Stage 1: Tool Discovery (What Tools Are Available?)
Before the LLM can choose a tool, it needs to know what tools exist.
In MCP, this happens through a process called tool listing.
When you connect an MCP server, it provides the LLM with a catalog of available tools. Each tool comes with:
- Name: A unique identifier (like "get_weather" or "send_email")
- Description: What the tool does in plain English
- Parameters: What inputs it needs (like location for weather, or recipient for email)
- Parameter descriptions: What each parameter means
Example: Weather Tool Definition
{
"name": "get_weather",
"description": "Get the current weather for any city in the world",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and country, e.g., 'Paris, France' or 'Tokyo, Japan'"
},
"unit": {
"type": "string",
"description": "Temperature unit: 'celsius' or 'fahrenheit'",
"enum": ["celsius", "fahrenheit"]
}
},
"required": ["location"]
}
}
Notice how detailed these descriptions are. They're not for humans — they're for the LLM to understand.
✅ Best Practice
Write tool descriptions as if you're explaining to a smart colleague.
Be specific about what the tool does and when to use it.
Example: Instead of "Gets weather" write "Get current weather conditions including temperature,
humidity, and forecast for any city worldwide."
Stage 2: Intent Recognition (What Does the User Want?)
When you ask the LLM a question, it first needs to understand your intent.
Let's see some examples:
User asks: "What's the weather like in London?"
LLM's internal reasoning:
- User wants current weather information
- Location mentioned: London
- No time unit specified (I'll assume current weather)
- Need to use external tool (I don't have real-time weather data)
User asks: "Tell me about the climate in London"
LLM's internal reasoning:
- User wants general climate information
- This is about long-term climate patterns, not current weather
- I can answer this from my training data
- No tool needed
See the difference? The LLM determines whether it needs a tool at all.
Stage 3: Tool Matching (Which Tool Fits Best?)
Once the LLM knows it needs a tool, it matches the user's request against available tools.
This is where things get interesting. Let's say your MCP server provides these tools:
- get_weather - Get current weather for a location
- get_forecast - Get 7-day weather forecast
- send_email - Send an email message
- search_database - Search company database
- create_calendar_event - Add event to calendar
User asks: "What will the weather be like in Paris next week?"
The LLM's matching process:
- ❌ send_email - Not about email
- ❌ search_database - Not about database
- ❌ create_calendar_event - Not about calendar
- ❌ get_weather - This is for current weather, but user wants forecast
- ✅ get_forecast - Perfect match! User wants future weather
The LLM compares the user's request against each tool's description and picks the best match.
Stage 4: Parameter Extraction (What Values to Pass?)
After choosing the tool, the LLM needs to figure out what parameters to send.
User's request: "What will the weather be like in Paris next week?"
Selected tool: get_forecast
Required parameters:
- location (required)
- days (optional, defaults to 7)
- unit (optional, defaults to celsius)
LLM's parameter extraction:
{
"location": "Paris, France",
"days": 7,
"unit": "celsius"
}
Notice how the LLM:
- Extracted "Paris" and expanded it to "Paris, France" for clarity
- Inferred "next week" means 7 days
- Used the default unit (celsius) since none was specified
💡 Key Insight
The LLM uses its language understanding to extract parameters from natural language. It can handle variations like "next week" (7 days), "tomorrow" (1 day), or "Fahrenheit" vs "F".
Stage 5: Structured Output Generation (The Function Call)
Finally, the LLM generates a structured output that tells your application what to do.
This output is typically in JSON format and looks like this:
{
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_forecast",
"arguments": "{\"location\": \"Paris, France\", \"days\": 7, \"unit\": \"celsius\"}"
}
}
]
}
Your application code then:
- Reads this JSON
- Calls the actual get_forecast function with those parameters
- Gets the real weather data from an API
- Sends the result back to the LLM
🎨 Real-World Example: Complete Flow
Let's walk through a complete example from start to finish.
Scenario: You're building a travel assistant with MCP tools.
Available MCP Tools:
- get_flight_prices - Search for flight prices
- book_flight - Book a specific flight
- get_hotel_availability - Check hotel rooms
- book_hotel - Reserve a hotel room
- get_weather - Get weather forecast
- create_itinerary - Create trip itinerary
User says: "I want to go to Barcelona next month. Find me cheap flights from London."
Step-by-Step Breakdown
Step 1: Intent Analysis
The LLM analyzes the request:
- User wants to travel from London to Barcelona
- Time frame: Next month
- Priority: Cheap flights (price-sensitive)
- Action needed: Search for flights (not booking yet)
Step 2: Tool Selection
The LLM scans available tools:
- ✅ get_flight_prices - This matches! User wants to search prices
- ❌ book_flight - User hasn't chosen a flight yet
- ❌ get_hotel_availability - Not about hotels
- ❌ book_hotel - Not about hotels
- ❌ get_weather - Not asking about weather
- ❌ create_itinerary - Too early for full itinerary
Selected tool: get_flight_prices
Step 3: Parameter Extraction
The tool requires these parameters:
{
"origin": "string (required) - Departure city",
"destination": "string (required) - Arrival city",
"departure_date": "string (required) - Format: YYYY-MM-DD",
"return_date": "string (optional) - Format: YYYY-MM-DD",
"max_price": "number (optional) - Maximum price in USD",
"sort_by": "string (optional) - 'price', 'duration', or 'rating'"
}
The LLM extracts:
{
"origin": "London, UK",
"destination": "Barcelona, Spain",
"departure_date": "2026-03-04", // Next month from today (Feb 4, 2026)
"sort_by": "price" // User said "cheap"
}
Note: The LLM calculated "next month" as March 4, 2026, and inferred "cheap" means sort by price.
Step 4: Generate Function Call
The LLM outputs:
{
"tool_calls": [
{
"id": "call_xyz789",
"type": "function",
"function": {
"name": "get_flight_prices",
"arguments": "{
\"origin\": \"London, UK\",
\"destination\": \"Barcelona, Spain\",
\"departure_date\": \"2026-03-04\",
\"sort_by\": \"price\"
}"
}
}
]
}
Step 5: Your Application Executes
Your code:
- Parses the JSON
- Calls your flight search API with those parameters
- Gets results (let's say 5 flight options)
- Sends results back to the LLM
Step 6: LLM Formats Final Response
The LLM receives the flight data and generates a user-friendly response:
"I found several affordable flights from London to Barcelona for early March: 1. Ryanair - £45 (Direct, 2h 15m) 2. EasyJet - £52 (Direct, 2h 20m) 3. British Airways - £89 (Direct, 2h 10m) The cheapest option is Ryanair at £45. Would you like me to check return flights or book one of these?"
That's the complete flow!
🧩 How the LLM Actually Chooses: The Matching Algorithm
You might be wondering: "How does the LLM match user requests to tools?"
The answer lies in semantic similarity — the LLM's core strength.
Semantic Matching Explained
LLMs understand language at a deep level. They can recognize that:
- "What's the weather?" ≈ "How's the climate?" ≈ "Is it raining?"
- "Book a flight" ≈ "Reserve plane tickets" ≈ "Get me on that flight"
- "Send an email" ≈ "Email this" ≈ "Write to them"
When matching tools, the LLM:
- Reads the tool's name and description
- Compares it semantically to the user's request
- Scores each tool based on relevance
- Picks the highest-scoring match
Example: Multiple Similar Tools
Imagine you have these tools:
1. send_email Description: "Send an email message to one or more recipients" 2. send_bulk_email Description: "Send the same email to multiple recipients at once (marketing)" 3. schedule_email Description: "Schedule an email to be sent at a specific future time"
User request: "Email our customers about the sale tomorrow at 9 AM"
LLM's matching logic:
-
send_email - Matches "email" but doesn't handle "tomorrow at 9 AM" scheduling
Score: 60% -
send_bulk_email - Matches "email our customers" (bulk) but doesn't handle scheduling
Score: 70% -
schedule_email - Matches "email" AND "tomorrow at 9 AM" (scheduled sending)
Score: 95%
Winner: schedule_email
The LLM recognized that both bulk sending AND scheduling were needed, and schedule_email likely supports both features.
✅ Pro Tip
When designing tools, make descriptions specific. Instead of "Sends emails" write "Send a single email to one recipient immediately." This helps the LLM distinguish between similar tools.
⚡ Advanced Concept: Multi-Tool Scenarios
Sometimes a single user request requires multiple tools.
Example: "Find the latest sales report and email it to my manager."
This requires two separate actions:
- Search for the sales report
- Email it to the manager
Sequential Tool Calling
The LLM handles this in sequence:
First call:
{
"function": {
"name": "search_files",
"arguments": "{\"query\": \"sales report\", \"type\": \"document\", \"sort\": \"date_desc\"}"
}
}
Your application searches and returns the file ID.
Second call:
{
"function": {
"name": "send_email",
"arguments": "{
\"recipient\": \"manager@company.com\",
\"subject\": \"Latest Sales Report\",
\"attachment_id\": \"file_12345\"
}"
}
}
Notice how the LLM:
- Used the result from the first call (file_12345) in the second call
- Broke down a complex request into logical steps
- Maintained context between calls
Parallel Tool Calling
Some LLMs can call multiple tools in parallel when they're independent.
Example: "What's the weather in London and Paris?"
These are independent requests, so the LLM can call both simultaneously:
{
"tool_calls": [
{
"id": "call_001",
"function": {
"name": "get_weather",
"arguments": "{\"location\": \"London, UK\"}"
}
},
{
"id": "call_002",
"function": {
"name": "get_weather",
"arguments": "{\"location\": \"Paris, France\"}"
}
}
]
}
This is faster because both calls execute at once instead of waiting for one to finish.
🎯 Optimizing Tool Selection: Semantic Tool Discovery
Here's a real problem: What if you have 100 tools?
Sending all 100 tool definitions to the LLM would:
- Consume tons of tokens (expensive!)
- Slow down response time
- Confuse the LLM with too many options
The solution? Semantic tool selection.
How Semantic Selection Works
Instead of sending all tools, we:
- Convert each tool description into a vector (mathematical representation)
- Store these vectors in a database
- When a user asks a question, convert their question to a vector
- Find the most similar tool vectors
- Send only the top 3-5 matching tools to the LLM
Real Example
Without semantic selection:
User: "Show me the README for my project" ↓ LLM receives: 93 tool definitions (47,000+ tokens) ↓ Cost: $0.50 per request Time: 5 seconds
With semantic selection:
User: "Show me the README for my project" ↓ Vector search finds top matches: 1. get_file_contents (95% match) 2. search_repository (82% match) 3. list_project_files (75% match) ↓ LLM receives: 3 tool definitions (1,500 tokens) ↓ Cost: $0.02 per request (96% cheaper!) Time: 1.5 seconds (70% faster!)
This technique reduced tokens by 89% while maintaining 100% accuracy.
✅ When to Use Semantic Selection
If you have more than 10-15 tools, semantic selection becomes essential. It dramatically reduces costs and improves performance without sacrificing accuracy.
🛠️ Practical Example: Building Your First MCP Tool
Let's create a simple calculator tool and see how the LLM chooses to use it.
Tool Definition
{
"name": "calculate",
"description": "Perform mathematical calculations. Can add, subtract, multiply, divide,
calculate percentages, square roots, and exponentials.",
"parameters": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "Mathematical expression to evaluate. Examples: '25 * 4',
'sqrt(144)', '(100 - 20) * 0.15'"
}
},
"required": ["expression"]
}
}
Test Cases
Test 1: Simple Math
User: "What's 234 times 56?"
LLM's decision:
- Recognizes this is a calculation
- Chooses the "calculate" tool
- Extracts expression: "234 * 56"
Function call:
{
"name": "calculate",
"arguments": "{\"expression\": \"234 * 56\"}"
}
Result: 13,104
Test 2: Complex Expression
User: "If I buy something for $150 and there's a 20% discount, what do I pay?"
LLM's decision:
- Recognizes this needs calculation
- Converts word problem to mathematical expression
- Original price: $150, Discount: 20%, Final price = 150 * (1 - 0.20)
Function call:
{
"name": "calculate",
"arguments": "{\"expression\": \"150 * (1 - 0.20)\"}"
}
Result: $120
Test 3: When NOT to Use the Tool
User: "What's a good tip percentage at restaurants?"
LLM's decision:
- This is not asking for a calculation
- It's asking for advice/knowledge
- NO tool needed
Direct answer: "A standard tip is typically 15-20% for good service..."
See how the LLM distinguished between calculation requests and knowledge questions?
🔒 Common Pitfalls and How to Avoid Them
Pitfall 1: Ambiguous Tool Descriptions
❌ Bad Example
{
"name": "search",
"description": "Search for stuff"
}
Problem: "Stuff" is too vague. Search what? Files? Emails? Web? Database?
✅ Good Example
{
"name": "search_company_files",
"description": "Search through company documents, PDFs, presentations, and spreadsheets
stored in the shared drive. Returns file names, locations, and preview snippets."
}
Clear, specific, and tells the LLM exactly what this tool does.
Pitfall 2: Overlapping Tools
Having tools that do similar things confuses the LLM.
❌ Confusing Setup
Tool 1: get_temperature - "Get the temperature" Tool 2: get_weather - "Get weather information" Tool 3: check_weather - "Check current weather"
The LLM will struggle to choose between these.
✅ Clear Setup
Tool 1: get_current_weather - "Get current temperature, humidity, wind speed,
and conditions for a specific location"
Tool 2: get_weather_forecast - "Get 7-day weather forecast including daily
high/low temperatures and precipitation chances"
Each tool has a distinct, clear purpose.
Pitfall 3: Missing Parameter Descriptions
❌ Insufficient
{
"parameters": {
"date": {
"type": "string"
}
}
}
What format? YYYY-MM-DD? MM/DD/YYYY? Timestamp? The LLM doesn't know.
✅ Complete
{
"parameters": {
"date": {
"type": "string",
"description": "Date in ISO 8601 format (YYYY-MM-DD). Example: '2026-03-15'",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
}
}
}
Crystal clear. The LLM knows exactly what format to use.
🎓 Advanced Topics: Fine-Tuning Tool Selection
Concept 1: Tool Choice Forcing
Sometimes you want to force the LLM to use a specific tool.
This is useful when:
- You're building a structured workflow
- The user explicitly requested a specific action
- You want deterministic behavior
Most LLM APIs allow you to specify:
{
"tool_choice": "auto" // LLM decides
}
or
{
"tool_choice": "required" // Must use a tool
}
or
{
"tool_choice": {
"type": "function",
"function": {"name": "specific_tool_name"} // Force specific tool
}
}
Concept 2: Context-Aware Selection
LLMs can use conversation history to make better tool choices.
Example conversation:
User: "Show me sales data from last quarter"
Assistant: *uses database_query tool* "Here's the Q4 2025 sales data..."
User: "Now calculate the growth rate"
The LLM remembers:
- Previous tool was database_query
- Data is already retrieved
- Now needs calculation on that data
- Should use calculate tool with the retrieved numbers
This context awareness leads to smarter tool selection.
Concept 3: Error Handling and Retries
What if a tool call fails?
Smart LLMs can:
- Try a different tool
- Adjust parameters and retry
- Ask the user for clarification
Example:
User: "Send email to john" LLM attempts: send_email(recipient: "john") ↓ Error: "Email address required, not just name" ↓ LLM tries: search_contacts(name: "john") ↓ Finds: john.smith@company.com ↓ LLM retries: send_email(recipient: "john.smith@company.com") ↓ Success!
The LLM adapted when the first approach failed.
📊 Measuring Tool Selection Success
How do you know if your tool descriptions are working well?
Track these metrics:
1. Tool Selection Accuracy
Measure: Is the LLM choosing the right tool?
Test with known scenarios:
- "Get weather" → Should use get_weather ✅
- "Book a flight" → Should use book_flight ✅
- "What's 2+2?" → Should use calculate ✅
Aim for 95%+ accuracy.
2. Parameter Extraction Quality
Measure: Are parameters correctly extracted?
- User: "Weather in NYC" → location: "New York City, NY" ✅
- User: "Next Friday" → date: "2026-02-14" ✅
- User: "In Fahrenheit" → unit: "fahrenheit" ✅
3. Tool Call Efficiency
Measure: How many tool calls needed?
- 1 call for simple tasks ✅
- 2-3 calls for complex multi-step tasks ✅
- 10+ calls for something simple ❌ (need better tool design)
Putting It All Together: Best Practices
✅ DO: Write Clear, Specific Descriptions
"Search employee database by name, department, or ID and return contact information" is better than "Search database."
✅ DO: Include Examples in Descriptions
"Date format: YYYY-MM-DD (example: '2026-03-15')" helps the LLM understand format requirements.
✅ DO: Make Tool Names Descriptive
"get_customer_order_history" is better than "get_data" or "fetch."
✅ DO: Group Related Parameters
Instead of 10 separate parameters, use objects: address: {street, city, state, zip}.
✅ DO: Test with Variations
Test with different phrasings: "weather," "temperature," "forecast," "climate" should all work.
❌ DON'T: Use Generic Names
"process," "handle," "manage" don't tell the LLM what the tool actually does.
❌ DON'T: Create Too Many Similar Tools
Combine related functionality into one tool with optional parameters instead of creating 5 variations.
❌ DON'T: Assume the LLM Knows Formats
Always specify: date formats, currency formats, enum values, min/max ranges.
❌ DON'T: Use Jargon Without Explanation
"Fetch BLOB from S3 bucket via ARN" might confuse the LLM. Explain: "Download file from Amazon S3 storage using its unique identifier."
🎯 Summary: The Complete Picture
Let's recap how LLMs choose which MCP tool to call:
- Tool Discovery: The MCP server provides a catalog of available tools with names, descriptions, and parameters.
- Intent Recognition: The LLM analyzes the user's request to understand what they want to accomplish.
- Semantic Matching: The LLM compares the request against each tool's description using natural language understanding.
- Parameter Extraction: The LLM extracts relevant information from the user's request to fill in the tool's parameters.
- Structured Output: The LLM generates a JSON function call that your application executes.
- Result Integration: The tool result is sent back to the LLM, which formats a final user-friendly response.
The key insight: LLMs don't magically know which tool to use. They rely on clear, well-written descriptions and semantic understanding to make smart choices.
Comments
Post a Comment