Ever wondered why the same AI sometimes gives you brilliant, creative responses and other times sounds robotic or nonsensical?
The secret lies in hyperparameters - the hidden control knobs that determine how your LLM thinks and responds!
Think of hyperparameters like the settings on your camera. Just as adjusting aperture, ISO, and shutter speed changes your photo completely, tweaking hyperparameters transforms how an LLM generates text.
By the end of this guide, you'll go from complete beginner to confidently controlling LLM behavior like a pro! 🎯
What Are Hyperparameters?
Let's start simple. When you ask ChatGPT or Claude a question, the model doesn't just "know" the answer.
Instead, it generates text word-by-word (technically, token-by-token). At each step, it looks at thousands of possible next words and picks one based on probability.
Hyperparameters are the settings that control HOW the model makes these choices. They're called "hyper" because they sit above the regular parameters (the model's learned weights) and guide the generation process.
Imagine you're at an ice cream shop with 50 flavors. Hyperparameters are like your decision-making rules:
- "Always pick chocolate" (low creativity/temperature)
- "Try something new every time" (high creativity/temperature)
- "Only consider the top 5 popular flavors" (top-k)
- "Pick from flavors that together make up 90% of people's choices" (top-p)
The Two Categories You Must Know 📚
Hyperparameters fall into two main groups:
1. Training Hyperparameters
These control how the model LEARNS during training. Think of them as settings for the student while studying.
Examples: learning rate, batch size, number of epochs, dropout rate.
Important: As an LLM engineer using APIs like OpenAI or Anthropic, you typically DON'T control these. The model is already trained!
2. Inference Hyperparameters (Our Focus!) ⭐
These control how the model GENERATES text when you use it. These are what you'll work with every day!
Examples: temperature, top-p, top-k, max tokens, frequency penalty.
Let's dive deep into each one with practical examples!
Temperature - The Creativity Dial 🌡️
What It Does
Temperature controls how random or creative the model's outputs are. It's the most important hyperparameter you'll use.
Range: Usually 0 to 2 (but 0 to 1 is most common)
How It Works Under the Hood
When the model predicts the next word, it assigns a probability to every word in its vocabulary (50,000+ words!).
Temperature adjusts these probabilities BEFORE the model picks:
- Lower temperature (0.1 - 0.5): Makes high-probability words even MORE likely, low-probability words even LESS likely
- Higher temperature (0.7 - 2.0): Flattens the distribution, giving unlikely words a better chance
- Factual accuracy (technical documentation)
- Code generation
- Mathematical calculations
- Consistent, predictable responses
- Translation tasks
- Creative writing (will be boring!)
- Brainstorming sessions
- Generating multiple diverse ideas
- Anything requiring variety
Practical Example - Same Prompt, Different Temperatures
Prompt: "Write a tagline for a coffee shop."
Temperature = 0.2 (Conservative):
"Fresh coffee, made daily." "Quality coffee at affordable prices." "Your neighborhood coffee shop."
Notice how these are safe, predictable, and similar to what you'd expect.
Temperature = 0.8 (Creative):
"Where every sip tells a story." "Brewing dreams, one cup at a time." "Life begins after coffee - start yours here."
More varied, creative, and engaging!
Temperature = 1.5 (Very High - Risky!):
"Velvet whispers dancing through ceramic portals of awakening." "The quantum foam of consciousness, espresso-style!"
Too creative - starting to lose coherence!
Start at temperature 0.7 as your default. Decrease if output is too random, increase if it's too boring. Rarely go above 1.0 unless you specifically want experimental outputs!
Code Example - Using Temperature
# Using OpenAI API
import openai
# Low temperature for factual task
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "What is 25 * 48?"}],
temperature=0.1 # Very deterministic
)
# High temperature for creative task
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Write a creative story opener"}],
temperature=0.9 # More creative
)
Top-P (Nucleus Sampling) - The Quality Filter 🎯
What It Does
While temperature changes HOW random selection is, top-p changes WHICH words can even be considered.
Range: 0 to 1 (most commonly 0.9 to 0.95)
How It Works
Top-p looks at all possible next words, ranked by probability, and draws a line where their cumulative probability reaches the p value.
Only words above that line are considered. Everything below gets ignored.
Let's visualize this with an example:
Prompt: "The cat sat on the ___"
Model's predictions (simplified):
- "mat" - 40% probability
- "floor" - 25% probability
- "chair" - 15% probability
- "roof" - 10% probability
- "table" - 5% probability
- "moon" - 3% probability
- "pizza" - 2% probability
With Top-P = 0.9:
The model adds up probabilities: 40% + 25% + 15% + 10% = 90%
It stops here! Only "mat", "floor", "chair", and "roof" are considered. "table", "moon", and "pizza" are excluded.
With Top-P = 0.6:
40% + 25% = 65% (exceeds 60%)
Only "mat" and "floor" are considered. Much more restrictive!
- You want creative variety
- Exploring different phrasings
- Generating multiple diverse outputs
- Storytelling or content creation
- Factual Q&A (use 0.1-0.3 instead)
- When you need the most accurate single answer
- Technical documentation requiring precision
Temperature vs Top-P - The Key Difference
This confuses everyone at first! Here's the clearest explanation:
- Temperature: Changes HOW MUCH randomness is in the selection among chosen words
- Top-P: Changes WHICH words are even allowed in the pool
The general recommendation from OpenAI and Anthropic is to adjust EITHER temperature OR top-p, but not both simultaneously!
If you're tuning temperature, leave top-p at 1.0. If you're tuning top-p, set temperature to 1.0.
Why? They interact in complex ways. Changing both makes it hard to predict behavior.
Code Example - Using Top-P
# Using Anthropic Claude API
import anthropic
client = anthropic.Anthropic(api_key="your-key")
# Focused, deterministic output
response = client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
top_p=0.3, # Very focused - only top 30% of probability mass
temperature=1.0, # Leave at default when using top_p
messages=[{
"role": "user",
"content": "Explain quantum entanglement"
}]
)
# Creative, diverse output
response = client.messages.create(
model="claude-3-sonnet-20240229",
max_tokens=1024,
top_p=0.95, # Allow more diversity
temperature=1.0,
messages=[{
"role": "user",
"content": "Write a fantasy story opening"
}]
)
Top-K - The Simple Limiter 🔢
What It Does
Top-k is simpler than top-p. It just says: "Only consider the k most likely words."
Range: Any positive integer (commonly 10 to 50)
How It Works
If top-k = 20, the model only looks at the 20 most probable next words. All others are ignored, regardless of their actual probabilities.
Example:
The model predicts 50,000 possible next words. With top-k = 30, it throws away 49,970 options and randomly picks from the remaining 30 based on their probabilities.
The Problem With Top-K
Top-k has a major flaw: it's inflexible.
Sometimes the model is very confident and the top 5 words capture 95% of probability. In this case, top-k = 50 unnecessarily includes 45 low-quality words!
Other times, the model is uncertain and even the top 50 words only capture 60% of probability. Here, top-k = 50 might cut off good options.
This is why top-p is generally preferred over top-k - it adapts to the model's confidence level dynamically.
Most commercial APIs (OpenAI, Anthropic) don't even expose top-k anymore. They use top-p instead. OpenAI API doesn't support top-k at all.
You'll mainly see top-k when using open-source models locally (via Hugging Face, vLLM, etc.).
Min-P
What It Does
Min-p is a breakthrough sampling method that's rapidly becoming the standard for open-source LLM deployments in 2025.
It was accepted as an oral presentation at ICLR 2025 (top 18 out of thousands of submissions)!
Range: 0.01 to 1.0 (commonly 0.05 to 0.1)
How It Works - The Smart Way
Min-p solves the biggest problem with top-p: temperature coupling.
Here's the issue with top-p: When you increase temperature to make output more creative, top-p accidentally lets in MORE words, even when the model is confident.
Min-p fixes this brilliantly:
- Find the probability of the most likely word
- Multiply it by min-p value (e.g., 0.1)
- Remove any word whose probability is below that threshold
Example:
Top word: "cat" with 60% probability
Min-p = 0.1
Threshold = 60% × 0.1 = 6%
Result: Any word with less than 6% probability gets excluded.
The Magic: When the model is confident (top word has high probability), min-p is strict and only keeps high-quality options.
When the model is uncertain (top word has low probability), min-p relaxes and allows more variety.
This adaptive behavior works beautifully at both low and high temperatures!
Why Min-P Is Taking Over
Research shows min-p consistently outperforms top-p, especially at higher temperatures.
It's now the default in:
- llama.cpp
- vLLM
- Hugging Face Transformers
- Ollama
- KoboldCpp
- Text Generation WebUI
Even cutting-edge models like DeepSeek-R1 recommend using min-p!
- For open-source/local models: Use temperature + min-p (0.05-0.1)
- For commercial APIs (OpenAI, Anthropic): Use temperature + top-p (they don't support min-p yet)
- For creative tasks: min-p = 0.05
- For general use: min-p = 0.1
Code Example - Using Min-P
# Using Hugging Face Transformers with Min-P
from transformers import AutoTokenizer, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3-8B")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3-8B")
prompt = "Write a creative story about a time traveler:"
inputs = tokenizer(prompt, return_tensors="pt")
# Generate with min-p
outputs = model.generate(
inputs["input_ids"],
do_sample=True,
temperature=0.8, # Moderate creativity
min_p=0.05, # Dynamic threshold based on confidence
max_length=200
)
print(tokenizer.decode(outputs[0]))
Max Tokens - Setting the Length Limit 📏
What It Does
Controls the maximum number of tokens (words/word pieces) the model can generate.
Important: 1 token ≈ 0.75 words in English. So 100 tokens ≈ 75 words.
Why It Matters
- Cost control: APIs charge per token. More tokens = more money!
- Response time: Shorter max_tokens = faster responses
- Preventing rambling: Stops the model from generating essays when you wanted a sentence
- Short answers: 50-150 tokens
- Paragraph responses: 200-500 tokens
- Long-form content: 1000-2000 tokens
- Articles/essays: 2000-4000 tokens
Code Example
# OpenAI API
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Explain photosynthesis briefly"}],
max_tokens=100 # Limits to roughly 75 words
)
# For longer responses
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Write a detailed blog post about AI"}],
max_tokens=2000 # Allows ~1500 words
)
Stop Sequences - Ending Generation Early 🛑
What It Does
Stop sequences are specific strings that tell the model: "Stop generating as soon as you see this."
Common Use Cases
Example 1: Generating a list
You want exactly 5 items in a list. Set stop sequence to "6." so generation stops after the 5th item.
Prompt: "List 5 benefits of exercise:" Stop sequence: ["6.", "\n\n"] Output: 1. Improves cardiovascular health 2. Builds muscle strength 3. Enhances mental clarity 4. Boosts immune system 5. Increases energy levels [Generation stops here instead of continuing]
Example 2: Q&A Format
Generating multiple Q&A pairs, stop after each answer:
Stop sequence: ["Q:", "Question:"] This ensures the model stops after completing one answer, not continuing to generate more questions.
Code Example
# Controlling output with stop sequences
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Generate a Python function and explain it:"
}],
stop=["# End of function", "\n\nNext"], # Stop at these markers
max_tokens=500
)
Frequency Penalty & Presence Penalty - Fighting Repetition 🔄
Frequency Penalty
Range: -2.0 to 2.0 (typically 0 to 1)
Penalizes tokens based on how often they've appeared in the generated text so far.
The more frequently a word appears, the less likely it becomes to appear again.
Use when: The model keeps repeating the same words or phrases
Presence Penalty
Range: -2.0 to 2.0 (typically 0 to 1)
Penalizes tokens that have appeared AT ALL in the generated text, regardless of frequency.
Even one occurrence makes that word less likely to appear again.
Use when: You want maximum diversity and to avoid any repetition
The Key Difference
- Frequency penalty: "This word appeared 5 times already? Make it way less likely."
- Presence penalty: "This word appeared even once? Make it somewhat less likely."
Use EITHER frequency penalty OR presence penalty, not both together.
Start with small values (0.1 to 0.5) and increase if repetition persists.
Practical Example
Without penalties (default):
"AI is transforming industries. AI is making processes efficient. AI is changing how we work. AI enables innovation..." (Repetitive! "AI" appears in every sentence)
With frequency_penalty = 0.5:
"Artificial intelligence is transforming industries. Machine learning makes processes efficient. These technologies change how we work. Innovation is enabled through automation..." (Much better! Uses synonyms and varied phrasing)
Code Example
# Avoiding repetitive content
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Write about the benefits of remote work"
}],
temperature=0.7,
frequency_penalty=0.3, # Penalize repeated words
presence_penalty=0.0, # Not using both simultaneously
max_tokens=300
)
# For maximum diversity
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Generate 10 unique product name ideas"
}],
temperature=0.8,
presence_penalty=0.6, # Strong penalty for any repetition
max_tokens=200
)
Context Window / Context Length - Memory Limits 🧠
What It Is
The context window is the total number of tokens the model can "see" at once - including your prompt AND the response.
Think of it as the model's short-term memory size.
Common Context Window Sizes (2025)
- GPT-4 Turbo: 128,000 tokens (~96,000 words)
- Claude 3.5 Sonnet: 200,000 tokens (~150,000 words)
- Gemini 1.5 Pro: 2 million tokens (~1.5 million words)
- Llama 3: 8,192 tokens (base) to 128,000 (extended versions)
Why It Matters
If your prompt + desired response exceeds the context window, the model will either:
- Truncate (cut off) the beginning of your prompt
- Refuse to process it
- Lose track of earlier information
Always calculate: Prompt tokens + Max tokens ≤ Context window
Example: If context window = 8,192 tokens and your prompt = 6,000 tokens, set max_tokens ≤ 2,192.
Practical Use Cases
Long documents: Claude 3.5 Sonnet can read entire books (200k tokens) and answer questions
Multi-turn conversations: The entire chat history needs to fit in the context window
Code repositories: Process multiple files at once for analysis
Putting It All Together - Real-World Scenarios 🎬
Scenario 1: Technical Documentation Generator
Goal: Generate accurate, consistent technical docs
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "system",
"content": "You are a technical writer creating API documentation."
}, {
"role": "user",
"content": "Document the /users/create endpoint"
}],
temperature=0.2, # Low - need consistency
top_p=1.0, # Not using top_p when using temperature
max_tokens=1000,
frequency_penalty=0.1, # Slight penalty to avoid repetitive phrasing
presence_penalty=0.0
)
Why these settings?
- Temperature 0.2: Ensures factual, consistent output
- Max tokens 1000: Enough for detailed docs but not excessive
- Small frequency penalty: Prevents monotonous language while staying accurate
Scenario 2: Creative Story Generator
Goal: Generate diverse, imaginative stories
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Write a sci-fi story about AI and humanity"
}],
temperature=0.9, # High - want creativity!
top_p=1.0,
max_tokens=2000, # Longer output for full story
frequency_penalty=0.3, # Avoid repetitive phrases
presence_penalty=0.0
)
Why these settings?
- Temperature 0.9: Encourages creative, unexpected plot twists
- Max tokens 2000: Room for a complete short story
- Frequency penalty 0.3: Prevents using same descriptive phrases repeatedly
Scenario 3: Chatbot for Customer Support
Goal: Helpful, professional, varied responses
response = anthropic.Client().messages.create(
model="claude-3-5-sonnet-20240620",
max_tokens=300,
temperature=1.0, # Default
top_p=0.7, # Focused but not too rigid
messages=[{
"role": "user",
"content": "How do I reset my password?"
}]
)
Why these settings?
- Top_p 0.7 with temp 1.0: Professional tone with some personality
- Max tokens 300: Concise answers that respect user's time
- No penalties: Allow natural, helpful language
Scenario 4: Code Generation
Goal: Syntactically correct, working code
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Write a Python function to merge two sorted lists"
}],
temperature=0.1, # Very low - need correctness
top_p=1.0,
max_tokens=500,
stop=["# End", "\n\n\n"], # Stop at code boundaries
frequency_penalty=0.0,
presence_penalty=0.0
)
Why these settings?
- Temperature 0.1: Minimal randomness, maximum correctness
- Stop sequences: Prevents generating extra unnecessary code
- No penalties: Let the model use standard coding patterns
Scenario 5: Brainstorming Session
Goal: Maximum diversity of ideas
# Generate multiple diverse ideas by making multiple calls
ideas = []
for i in range(5): # Get 5 different idea sets
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{
"role": "user",
"content": "Suggest innovative mobile app ideas for fitness"
}],
temperature=1.0, # High creativity
top_p=0.95, # Allow diverse tokens
max_tokens=200,
presence_penalty=0.7, # High - want unique ideas each time!
n=1
)
ideas.append(response.choices[0].message.content)
Why these settings?
- Temperature 1.0: Good creative range
- High presence penalty: Forces the model to explore different concepts each time
- Multiple calls: Generates truly diverse idea sets
Advanced Topic: Sampling Methods Comparison 📊
Greedy Decoding vs Sampling
When temperature = 0 (or very close), models use greedy decoding: always pick the most probable next token.
Result: Completely deterministic. Same input = same output every time.
When temperature > 0, models use sampling: randomly pick from probability distribution.
Result: Different outputs each time (stochastic).
Which Sampling Method Should You Use?
Here's the decision tree for 2025:
Using Commercial API (OpenAI, Anthropic, Google)?
- → Use Temperature + Top-P
- → Set one to default (1.0), tune the other
Using Open-Source/Local Model?
- → Use Temperature + Min-P (0.05-0.1)
- → Avoid Top-P and Top-K
Need Absolute Determinism?
- → Set Temperature = 0 (or use greedy decoding)
- → Also consider setting a random seed for reproducibility
Common Mistakes to Avoid ⚠️
Why it's bad: They interact unpredictably. You won't know which is causing behavior changes.
Fix: Pick one to tune. Set the other to 1.0.
Why it's bad: Introduces randomness where you need accuracy. Model might generate plausible-sounding but incorrect facts.
Fix: Use temperature ≤ 0.3 for any factual/technical content.
Why it's bad: Response gets cut off mid-sentence. Looks unprofessional and confusing.
Fix: Calculate expected response length + 20% buffer. Better to have room than to truncate!
Why it's bad: Model silently truncates your prompt or errors out. You lose important context.
Fix: Always check: prompt_tokens + max_tokens ≤ context_window. Count tokens before sending!
Why it's bad: Setting frequency/presence penalty too high (>1.0) makes text sound unnatural and forced.
Fix: Start small (0.1-0.3). Increase gradually only if needed. Rarely exceed 0.6.
Why it's bad: You stick with defaults that might not be optimal for your specific use case.
Fix: Run A/B tests. Generate same prompt with different hyperparameters. Compare results.
Testing and Optimization Workflow 🔬
Here's a systematic approach to finding the best hyperparameters for your use case:
Step 1: Start with Baseline
Baseline Settings: - Temperature: 0.7 - Top-P: 1.0 (or Min-P: 0.1 for open-source) - Max tokens: Reasonable for your task - Penalties: 0.0
Step 2: Identify Your Priority
Ask yourself: What matters most?
- Accuracy/Correctness? → Lower temperature (0.1-0.3)
- Creativity/Diversity? → Higher temperature (0.8-1.0)
- Avoiding repetition? → Add frequency penalty (0.2-0.5)
- Speed? → Reduce max_tokens
- Cost? → Reduce max_tokens, optimize prompt length
Step 3: Adjust One Parameter at a Time
This is crucial! If you change multiple things, you won't know what helped.
Test 1: Baseline (temp=0.7) Test 2: temp=0.3 Test 3: temp=0.5 Test 4: temp=0.9 Pick the best temperature, then move to next parameter. Test 5: Best temp + frequency_penalty=0.2 Test 6: Best temp + frequency_penalty=0.4 ...and so on
Step 4: Create a Test Suite
# Systematic hyperparameter testing
import openai
test_prompts = [
"Explain quantum computing",
"Write a creative story opening",
"Debug this Python code: [code]"
]
temperatures = [0.2, 0.5, 0.7, 1.0]
results = {}
for temp in temperatures:
results[temp] = []
for prompt in test_prompts:
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": prompt}],
temperature=temp,
max_tokens=200
)
results[temp].append(response.choices[0].message.content)
# Now compare results[0.2] vs results[0.5] etc.
# Choose the temperature that gives best results across all prompts
Step 5: Measure Quality
Don't just eyeball results. Use metrics:
- For factual accuracy: Compare against ground truth, count errors
- For creativity: Measure diversity (unique words, sentence structures)
- For user satisfaction: Run A/B tests with real users
- For code: Run unit tests, measure syntax errors
Hyperparameter Cheat Sheet 📋
| Parameter | Range | Use Low For | Use High For |
|---|---|---|---|
| Temperature | 0-2 |
Factual Q&A Code generation Technical docs Translations |
Creative writing Brainstorming Diverse outputs Storytelling |
| Top-P | 0-1 |
Precise answers Focused output Single best response |
Exploring options Creative tasks Diverse vocabulary |
| Min-P | 0.01-1 |
General purpose (0.1) Focused output |
Not recommended (Use 0.05-0.1) |
| Top-K | 1-∞ |
Very focused (5-10) Deterministic |
More variety (40-50) (But use Top-P instead!) |
| Max Tokens | 1-∞ |
Quick answers Chat responses Cost control |
Long-form content Articles Detailed analysis |
| Frequency Penalty | -2 to 2 |
Technical writing (Need specific terms) |
Avoiding repetition Varied language Creative content |
| Presence Penalty | -2 to 2 |
Allow key term usage Domain-specific text |
Maximum diversity Exploring topics Idea generation |
Quick Reference - Common Tasks 🎯
📝 Writing Blog Posts
temperature: 0.7 top_p: 1.0 max_tokens: 2000 frequency_penalty: 0.3 presence_penalty: 0.0
💻 Code Generation
temperature: 0.2 top_p: 1.0 max_tokens: 1000 frequency_penalty: 0.0 presence_penalty: 0.0
🎨 Creative Fiction
temperature: 0.9 top_p: 1.0 max_tokens: 3000 frequency_penalty: 0.3 presence_penalty: 0.1
❓ Q&A / Customer Support
temperature: 1.0 top_p: 0.7 max_tokens: 300 frequency_penalty: 0.0 presence_penalty: 0.0
📊 Data Analysis Reports
temperature: 0.3 top_p: 1.0 max_tokens: 1500 frequency_penalty: 0.1 presence_penalty: 0.0
💡 Brainstorming Ideas
temperature: 1.0 top_p: 0.95 max_tokens: 500 frequency_penalty: 0.0 presence_penalty: 0.6
Advanced Techniques for Experts 🚀
1. Adaptive Hyperparameter Selection
Instead of using fixed values, adjust hyperparameters based on the task:
def get_optimal_params(task_type):
"""Return hyperparameters optimized for task type"""
configs = {
"factual": {
"temperature": 0.2,
"top_p": 1.0,
"max_tokens": 500,
"frequency_penalty": 0.0
},
"creative": {
"temperature": 0.9,
"top_p": 1.0,
"max_tokens": 2000,
"frequency_penalty": 0.3
},
"code": {
"temperature": 0.1,
"top_p": 1.0,
"max_tokens": 1500,
"frequency_penalty": 0.0
},
"conversational": {
"temperature": 1.0,
"top_p": 0.8,
"max_tokens": 300,
"frequency_penalty": 0.0
}
}
return configs.get(task_type, configs["conversational"])
# Usage
params = get_optimal_params("creative")
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": "Write a story"}],
**params # Unpack the optimal parameters
)
2. Temperature Scheduling
For long-form generation, start high (creative) and gradually decrease (focused):
def generate_with_temperature_schedule(prompt, stages):
"""
Generate text with decreasing temperature across stages
stages: [(temperature, max_tokens), ...]
"""
full_response = ""
conversation = [{"role": "user", "content": prompt}]
for temp, tokens in stages:
response = openai.ChatCompletion.create(
model="gpt-4",
messages=conversation,
temperature=temp,
max_tokens=tokens
)
chunk = response.choices[0].message.content
full_response += chunk
# Add assistant's response to conversation
conversation.append({"role": "assistant", "content": chunk})
conversation.append({
"role": "user",
"content": "Continue writing..."
})
return full_response
# Example: Start creative, end focused
result = generate_with_temperature_schedule(
"Write a research paper about AI safety",
stages=[
(0.9, 500), # Creative introduction
(0.7, 1000), # Detailed body
(0.3, 500) # Precise conclusion
]
)
3. Multi-Sample Voting
Generate multiple responses with high temperature, then use a deterministic model to pick the best:
def generate_with_voting(prompt, n_samples=5):
"""Generate multiple samples and pick the best via voting"""
# Step 1: Generate diverse candidates
candidates = []
for _ in range(n_samples):
response = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": prompt}],
temperature=0.9, # High diversity
max_tokens=200
)
candidates.append(response.choices[0].message.content)
# Step 2: Use low-temp model to pick best
voting_prompt = f"""
Here are {n_samples} candidate responses to: "{prompt}"
Candidates:
{chr(10).join(f"{i+1}. {c}" for i, c in enumerate(candidates))}
Which response is the best? Reply with just the number (1-{n_samples}).
"""
vote = openai.ChatCompletion.create(
model="gpt-4",
messages=[{"role": "user", "content": voting_prompt}],
temperature=0.1, # Deterministic voting
max_tokens=10
)
best_idx = int(vote.choices[0].message.content.strip()) - 1
return candidates[best_idx]
# Usage
best_response = generate_with_voting(
"Explain quantum entanglement simply"
)
4. Context-Aware Max Tokens
Dynamically calculate max_tokens based on remaining context window:
import tiktoken # OpenAI's tokenizer
def count_tokens(text, model="gpt-4"):
"""Count tokens in a text string"""
encoding = tiktoken.encoding_for_model(model)
return len(encoding.encode(text))
def smart_generate(messages, model="gpt-4", context_window=8192):
"""Automatically set max_tokens based on available context"""
# Count tokens in messages
prompt_tokens = sum(
count_tokens(msg["content"], model)
for msg in messages
)
# Calculate safe max_tokens
# Leave 10% buffer for safety
available = context_window - prompt_tokens
safe_max_tokens = int(available * 0.9)
if safe_max_tokens < 100:
raise ValueError("Not enough context window remaining!")
response = openai.ChatCompletion.create(
model=model,
messages=messages,
max_tokens=min(safe_max_tokens, 4000), # Cap at 4000
temperature=0.7
)
return response
# Usage - automatically handles token limits
response = smart_generate([
{"role": "user", "content": "Your very long prompt here..."}
])
Troubleshooting Guide 🔧
Problem: Output is Too Random/Nonsensical
Symptoms: Text doesn't make sense, random words, incoherent sentences
Solutions:
- Reduce temperature from 1.0 → 0.7 → 0.5
- Reduce top_p from 0.95 → 0.8 → 0.6
- If using top-k, reduce from 50 → 30 → 10
- Check if you're accidentally using temperature > 1.5
Problem: Output is Too Repetitive/Boring
Symptoms: Same phrases over and over, monotonous language, always says the same thing
Solutions:
- Increase temperature from 0.5 → 0.7 → 0.9
- Add frequency_penalty: 0.2 → 0.4 → 0.6
- Try presence_penalty if frequency_penalty doesn't help
- Increase top_p to allow more token diversity
Problem: Responses Cut Off Mid-Sentence
Symptoms: Output ends abruptly, incomplete thoughts
Solutions:
- Increase max_tokens (double it as a test)
- Check you're not hitting context window limit
- Remove or adjust stop sequences if using them
- Calculate: prompt_tokens + max_tokens must fit in context_window
Problem: Output is Factually Incorrect
Symptoms: Plausible-sounding but wrong information, hallucinations
Solutions:
- Reduce temperature to ≤ 0.3 for factual tasks
- Use top_p around 0.1-0.3 instead of high values
- Add "Be factual and accurate" to system prompt
- Consider using retrieval-augmented generation (RAG) for better factual grounding
Problem: Responses Too Expensive
Symptoms: API costs are high
Solutions:
- Reduce max_tokens to minimum needed
- Optimize your prompts to be more concise
- Use a smaller model (e.g., GPT-3.5 instead of GPT-4) where appropriate
- Cache common responses
- Implement rate limiting
The Future of Hyperparameters (2025 and Beyond) 🔮
Emerging Trends
1. Min-P Becoming Standard
Expect commercial APIs (OpenAI, Anthropic) to add Min-P support soon. It's proven superior to Top-P in research.
2. Reasoning Models with Locked Parameters
Advanced reasoning models (like o1, o3) often lock their hyperparameters. You can't change temperature or top-p because they're optimized internally.
3. Automatic Hyperparameter Tuning
ML systems that learn optimal hyperparameters for your specific use case by analyzing your feedback.
4. Context-Aware Adaptive Sampling
Models that automatically adjust sampling strategy based on what they're generating (creative vs factual) without manual tuning.
Research to Watch
- Top-n-sigma: Temperature-invariant sampling (published ACL 2025)
- Mirostat: Adaptive sampling based on perplexity control
- Dynamic Temperature: Automatically adjust temperature per-token
- Beam Search Improvements: Better quality-diversity tradeoffs
Summary - Your Hyperparameter Journey 🎓
Congratulations! You've gone from zero knowledge to understanding hyperparameters like a pro!
Let's recap the essentials:
- Temperature controls creativity (0.2 = factual, 0.9 = creative)
- Top-P filters which words are considered (0.1 = focused, 0.95 = diverse)
- Min-P is the new superior alternative to Top-P for open-source models
- Max Tokens limits response length - plan it based on context window
- Penalties fight repetition - use sparingly (0.1-0.5 range)
- Tune ONE parameter at a time - never temperature AND top-p together
- Different tasks need different settings - no one-size-fits-all
Document your best configurations! Create a personal "cheat sheet" of what works for your specific use cases. This saves huge amounts of time later.
Further Learning Resources 📚
Want to dive deeper? Check out these resources:
Research Papers
- "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs" (Nguyen et al., ICLR 2025)
- "The Curious Case of Neural Text Degeneration" (Holtzman et al., 2020) - Introduced Top-P
- "Top-n-sigma: Temperature-Invariant Sampling" (Tang et al., ACL 2025)
Documentation
- OpenAI API Parameters: https://platform.openai.com/docs/api-reference/chat
- Anthropic Claude Parameters: https://docs.anthropic.com/claude/reference
- Hugging Face Generation: https://huggingface.co/docs/transformers/main_classes/text_generation
Conclusion 🎉
Hyperparameters are the secret sauce that transforms generic LLM outputs into exactly what you need.
Happy LLM Enginneering ! 🤖
Comments
Post a Comment