Skip to main content

Mastering LLM Hyperparameters

Calculating read time…

Ever wondered why the same AI sometimes gives you brilliant, creative responses and other times sounds robotic or nonsensical?

The secret lies in hyperparameters - the hidden control knobs that determine how your LLM thinks and responds!

Think of hyperparameters like the settings on your camera. Just as adjusting aperture, ISO, and shutter speed changes your photo completely, tweaking hyperparameters transforms how an LLM generates text.

By the end of this guide, you'll go from complete beginner to confidently controlling LLM behavior like a pro! 🎯

What Are Hyperparameters?

Let's start simple. When you ask ChatGPT or Claude a question, the model doesn't just "know" the answer.

Instead, it generates text word-by-word (technically, token-by-token). At each step, it looks at thousands of possible next words and picks one based on probability.

Hyperparameters are the settings that control HOW the model makes these choices. They're called "hyper" because they sit above the regular parameters (the model's learned weights) and guide the generation process.

💡 Real-World Analogy:

Imagine you're at an ice cream shop with 50 flavors. Hyperparameters are like your decision-making rules:

  • "Always pick chocolate" (low creativity/temperature)
  • "Try something new every time" (high creativity/temperature)
  • "Only consider the top 5 popular flavors" (top-k)
  • "Pick from flavors that together make up 90% of people's choices" (top-p)

The Two Categories You Must Know 📚

Hyperparameters fall into two main groups:

1. Training Hyperparameters

These control how the model LEARNS during training. Think of them as settings for the student while studying.

Examples: learning rate, batch size, number of epochs, dropout rate.

Important: As an LLM engineer using APIs like OpenAI or Anthropic, you typically DON'T control these. The model is already trained!

2. Inference Hyperparameters (Our Focus!) ⭐

These control how the model GENERATES text when you use it. These are what you'll work with every day!

Examples: temperature, top-p, top-k, max tokens, frequency penalty.

Let's dive deep into each one with practical examples!

Temperature - The Creativity Dial 🌡️

What It Does

Temperature controls how random or creative the model's outputs are. It's the most important hyperparameter you'll use.

Range: Usually 0 to 2 (but 0 to 1 is most common)

How It Works Under the Hood

When the model predicts the next word, it assigns a probability to every word in its vocabulary (50,000+ words!).

Temperature adjusts these probabilities BEFORE the model picks:

  • Lower temperature (0.1 - 0.5): Makes high-probability words even MORE likely, low-probability words even LESS likely
  • Higher temperature (0.7 - 2.0): Flattens the distribution, giving unlikely words a better chance
✅ DO: Use Low Temperature (0.1 - 0.3) When You Need:
  • Factual accuracy (technical documentation)
  • Code generation
  • Mathematical calculations
  • Consistent, predictable responses
  • Translation tasks
❌ DON'T: Use Low Temperature For:
  • Creative writing (will be boring!)
  • Brainstorming sessions
  • Generating multiple diverse ideas
  • Anything requiring variety

Practical Example - Same Prompt, Different Temperatures

Prompt: "Write a tagline for a coffee shop."

Temperature = 0.2 (Conservative):

"Fresh coffee, made daily."
"Quality coffee at affordable prices."
"Your neighborhood coffee shop."

Notice how these are safe, predictable, and similar to what you'd expect.

Temperature = 0.8 (Creative):

"Where every sip tells a story."
"Brewing dreams, one cup at a time."
"Life begins after coffee - start yours here."

More varied, creative, and engaging!

Temperature = 1.5 (Very High - Risky!):

"Velvet whispers dancing through ceramic portals of awakening."
"The quantum foam of consciousness, espresso-style!"

Too creative - starting to lose coherence!

💡 Pro Tip:

Start at temperature 0.7 as your default. Decrease if output is too random, increase if it's too boring. Rarely go above 1.0 unless you specifically want experimental outputs!

Code Example - Using Temperature


# Using OpenAI API
import openai

# Low temperature for factual task
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "What is 25 * 48?"}],
    temperature=0.1  # Very deterministic
)

# High temperature for creative task
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Write a creative story opener"}],
    temperature=0.9  # More creative
)

Top-P (Nucleus Sampling) - The Quality Filter 🎯

What It Does

While temperature changes HOW random selection is, top-p changes WHICH words can even be considered.

Range: 0 to 1 (most commonly 0.9 to 0.95)

How It Works

Top-p looks at all possible next words, ranked by probability, and draws a line where their cumulative probability reaches the p value.

Only words above that line are considered. Everything below gets ignored.

Let's visualize this with an example:

Prompt: "The cat sat on the ___"

Model's predictions (simplified):

  • "mat" - 40% probability
  • "floor" - 25% probability
  • "chair" - 15% probability
  • "roof" - 10% probability
  • "table" - 5% probability
  • "moon" - 3% probability
  • "pizza" - 2% probability

With Top-P = 0.9:

The model adds up probabilities: 40% + 25% + 15% + 10% = 90%

It stops here! Only "mat", "floor", "chair", and "roof" are considered. "table", "moon", and "pizza" are excluded.

With Top-P = 0.6:

40% + 25% = 65% (exceeds 60%)

Only "mat" and "floor" are considered. Much more restrictive!

✅ DO: Use High Top-P (0.9 - 0.95) When:
  • You want creative variety
  • Exploring different phrasings
  • Generating multiple diverse outputs
  • Storytelling or content creation
❌ DON'T: Use High Top-P For:
  • Factual Q&A (use 0.1-0.3 instead)
  • When you need the most accurate single answer
  • Technical documentation requiring precision

Temperature vs Top-P - The Key Difference

This confuses everyone at first! Here's the clearest explanation:

  • Temperature: Changes HOW MUCH randomness is in the selection among chosen words
  • Top-P: Changes WHICH words are even allowed in the pool
💡 Golden Rule:

The general recommendation from OpenAI and Anthropic is to adjust EITHER temperature OR top-p, but not both simultaneously!

If you're tuning temperature, leave top-p at 1.0. If you're tuning top-p, set temperature to 1.0.

Why? They interact in complex ways. Changing both makes it hard to predict behavior.

Code Example - Using Top-P


# Using Anthropic Claude API
import anthropic

client = anthropic.Anthropic(api_key="your-key")

# Focused, deterministic output
response = client.messages.create(
    model="claude-3-sonnet-20240229",
    max_tokens=1024,
    top_p=0.3,  # Very focused - only top 30% of probability mass
    temperature=1.0,  # Leave at default when using top_p
    messages=[{
        "role": "user", 
        "content": "Explain quantum entanglement"
    }]
)

# Creative, diverse output
response = client.messages.create(
    model="claude-3-sonnet-20240229",
    max_tokens=1024,
    top_p=0.95,  # Allow more diversity
    temperature=1.0,
    messages=[{
        "role": "user", 
        "content": "Write a fantasy story opening"
    }]
)

Top-K - The Simple Limiter 🔢

What It Does

Top-k is simpler than top-p. It just says: "Only consider the k most likely words."

Range: Any positive integer (commonly 10 to 50)

How It Works

If top-k = 20, the model only looks at the 20 most probable next words. All others are ignored, regardless of their actual probabilities.

Example:

The model predicts 50,000 possible next words. With top-k = 30, it throws away 49,970 options and randomly picks from the remaining 30 based on their probabilities.

The Problem With Top-K

Top-k has a major flaw: it's inflexible.

Sometimes the model is very confident and the top 5 words capture 95% of probability. In this case, top-k = 50 unnecessarily includes 45 low-quality words!

Other times, the model is uncertain and even the top 50 words only capture 60% of probability. Here, top-k = 50 might cut off good options.

This is why top-p is generally preferred over top-k - it adapts to the model's confidence level dynamically.

💡 Modern Practice:

Most commercial APIs (OpenAI, Anthropic) don't even expose top-k anymore. They use top-p instead. OpenAI API doesn't support top-k at all.

You'll mainly see top-k when using open-source models locally (via Hugging Face, vLLM, etc.).

Min-P

What It Does

Min-p is a breakthrough sampling method that's rapidly becoming the standard for open-source LLM deployments in 2025.

It was accepted as an oral presentation at ICLR 2025 (top 18 out of thousands of submissions)!

Range: 0.01 to 1.0 (commonly 0.05 to 0.1)

How It Works - The Smart Way

Min-p solves the biggest problem with top-p: temperature coupling.

Here's the issue with top-p: When you increase temperature to make output more creative, top-p accidentally lets in MORE words, even when the model is confident.

Min-p fixes this brilliantly:

  1. Find the probability of the most likely word
  2. Multiply it by min-p value (e.g., 0.1)
  3. Remove any word whose probability is below that threshold

Example:

Top word: "cat" with 60% probability

Min-p = 0.1

Threshold = 60% × 0.1 = 6%

Result: Any word with less than 6% probability gets excluded.

The Magic: When the model is confident (top word has high probability), min-p is strict and only keeps high-quality options.

When the model is uncertain (top word has low probability), min-p relaxes and allows more variety.

This adaptive behavior works beautifully at both low and high temperatures!

Why Min-P Is Taking Over

Research shows min-p consistently outperforms top-p, especially at higher temperatures.

It's now the default in:

  • llama.cpp
  • vLLM
  • Hugging Face Transformers
  • Ollama
  • KoboldCpp
  • Text Generation WebUI

Even cutting-edge models like DeepSeek-R1 recommend using min-p!

✅ Best Practice for 2025:
  • For open-source/local models: Use temperature + min-p (0.05-0.1)
  • For commercial APIs (OpenAI, Anthropic): Use temperature + top-p (they don't support min-p yet)
  • For creative tasks: min-p = 0.05
  • For general use: min-p = 0.1

Code Example - Using Min-P


# Using Hugging Face Transformers with Min-P
from transformers import AutoTokenizer, AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3-8B")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3-8B")

prompt = "Write a creative story about a time traveler:"
inputs = tokenizer(prompt, return_tensors="pt")

# Generate with min-p
outputs = model.generate(
    inputs["input_ids"],
    do_sample=True,
    temperature=0.8,      # Moderate creativity
    min_p=0.05,          # Dynamic threshold based on confidence
    max_length=200
)

print(tokenizer.decode(outputs[0]))

Max Tokens - Setting the Length Limit 📏

What It Does

Controls the maximum number of tokens (words/word pieces) the model can generate.

Important: 1 token ≈ 0.75 words in English. So 100 tokens ≈ 75 words.

Why It Matters

  • Cost control: APIs charge per token. More tokens = more money!
  • Response time: Shorter max_tokens = faster responses
  • Preventing rambling: Stops the model from generating essays when you wanted a sentence
✅ Recommended Max Token Values:
  • Short answers: 50-150 tokens
  • Paragraph responses: 200-500 tokens
  • Long-form content: 1000-2000 tokens
  • Articles/essays: 2000-4000 tokens

Code Example


# OpenAI API
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Explain photosynthesis briefly"}],
    max_tokens=100  # Limits to roughly 75 words
)

# For longer responses
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Write a detailed blog post about AI"}],
    max_tokens=2000  # Allows ~1500 words
)

Stop Sequences - Ending Generation Early 🛑

What It Does

Stop sequences are specific strings that tell the model: "Stop generating as soon as you see this."

Common Use Cases

Example 1: Generating a list

You want exactly 5 items in a list. Set stop sequence to "6." so generation stops after the 5th item.

Prompt: "List 5 benefits of exercise:"

Stop sequence: ["6.", "\n\n"]

Output:
1. Improves cardiovascular health
2. Builds muscle strength
3. Enhances mental clarity
4. Boosts immune system
5. Increases energy levels

[Generation stops here instead of continuing]

Example 2: Q&A Format

Generating multiple Q&A pairs, stop after each answer:

Stop sequence: ["Q:", "Question:"]

This ensures the model stops after completing one answer, 
not continuing to generate more questions.

Code Example


# Controlling output with stop sequences
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "user", 
        "content": "Generate a Python function and explain it:"
    }],
    stop=["# End of function", "\n\nNext"],  # Stop at these markers
    max_tokens=500
)

Frequency Penalty & Presence Penalty - Fighting Repetition 🔄

Frequency Penalty

Range: -2.0 to 2.0 (typically 0 to 1)

Penalizes tokens based on how often they've appeared in the generated text so far.

The more frequently a word appears, the less likely it becomes to appear again.

Use when: The model keeps repeating the same words or phrases

Presence Penalty

Range: -2.0 to 2.0 (typically 0 to 1)

Penalizes tokens that have appeared AT ALL in the generated text, regardless of frequency.

Even one occurrence makes that word less likely to appear again.

Use when: You want maximum diversity and to avoid any repetition

The Key Difference

  • Frequency penalty: "This word appeared 5 times already? Make it way less likely."
  • Presence penalty: "This word appeared even once? Make it somewhat less likely."
💡 Best Practice:

Use EITHER frequency penalty OR presence penalty, not both together.

Start with small values (0.1 to 0.5) and increase if repetition persists.

Practical Example

Without penalties (default):

"AI is transforming industries. AI is making processes efficient. 
AI is changing how we work. AI enables innovation..."

(Repetitive! "AI" appears in every sentence)

With frequency_penalty = 0.5:

"Artificial intelligence is transforming industries. 
Machine learning makes processes efficient. 
These technologies change how we work. 
Innovation is enabled through automation..."

(Much better! Uses synonyms and varied phrasing)

Code Example


# Avoiding repetitive content
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "user", 
        "content": "Write about the benefits of remote work"
    }],
    temperature=0.7,
    frequency_penalty=0.3,  # Penalize repeated words
    presence_penalty=0.0,   # Not using both simultaneously
    max_tokens=300
)

# For maximum diversity
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "user", 
        "content": "Generate 10 unique product name ideas"
    }],
    temperature=0.8,
    presence_penalty=0.6,   # Strong penalty for any repetition
    max_tokens=200
)

Context Window / Context Length - Memory Limits 🧠

What It Is

The context window is the total number of tokens the model can "see" at once - including your prompt AND the response.

Think of it as the model's short-term memory size.

Common Context Window Sizes (2025)

  • GPT-4 Turbo: 128,000 tokens (~96,000 words)
  • Claude 3.5 Sonnet: 200,000 tokens (~150,000 words)
  • Gemini 1.5 Pro: 2 million tokens (~1.5 million words)
  • Llama 3: 8,192 tokens (base) to 128,000 (extended versions)

Why It Matters

If your prompt + desired response exceeds the context window, the model will either:

  1. Truncate (cut off) the beginning of your prompt
  2. Refuse to process it
  3. Lose track of earlier information
💡 Planning Tip:

Always calculate: Prompt tokens + Max tokens ≤ Context window

Example: If context window = 8,192 tokens and your prompt = 6,000 tokens, set max_tokens ≤ 2,192.

Practical Use Cases

Long documents: Claude 3.5 Sonnet can read entire books (200k tokens) and answer questions

Multi-turn conversations: The entire chat history needs to fit in the context window

Code repositories: Process multiple files at once for analysis

Putting It All Together - Real-World Scenarios 🎬

Scenario 1: Technical Documentation Generator

Goal: Generate accurate, consistent technical docs


response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "system",
        "content": "You are a technical writer creating API documentation."
    }, {
        "role": "user",
        "content": "Document the /users/create endpoint"
    }],
    temperature=0.2,        # Low - need consistency
    top_p=1.0,             # Not using top_p when using temperature
    max_tokens=1000,
    frequency_penalty=0.1,  # Slight penalty to avoid repetitive phrasing
    presence_penalty=0.0
)

Why these settings?

  • Temperature 0.2: Ensures factual, consistent output
  • Max tokens 1000: Enough for detailed docs but not excessive
  • Small frequency penalty: Prevents monotonous language while staying accurate

Scenario 2: Creative Story Generator

Goal: Generate diverse, imaginative stories


response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "user",
        "content": "Write a sci-fi story about AI and humanity"
    }],
    temperature=0.9,        # High - want creativity!
    top_p=1.0,
    max_tokens=2000,       # Longer output for full story
    frequency_penalty=0.3,  # Avoid repetitive phrases
    presence_penalty=0.0
)

Why these settings?

  • Temperature 0.9: Encourages creative, unexpected plot twists
  • Max tokens 2000: Room for a complete short story
  • Frequency penalty 0.3: Prevents using same descriptive phrases repeatedly

Scenario 3: Chatbot for Customer Support

Goal: Helpful, professional, varied responses


response = anthropic.Client().messages.create(
    model="claude-3-5-sonnet-20240620",
    max_tokens=300,
    temperature=1.0,       # Default
    top_p=0.7,            # Focused but not too rigid
    messages=[{
        "role": "user",
        "content": "How do I reset my password?"
    }]
)

Why these settings?

  • Top_p 0.7 with temp 1.0: Professional tone with some personality
  • Max tokens 300: Concise answers that respect user's time
  • No penalties: Allow natural, helpful language

Scenario 4: Code Generation

Goal: Syntactically correct, working code


response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{
        "role": "user",
        "content": "Write a Python function to merge two sorted lists"
    }],
    temperature=0.1,       # Very low - need correctness
    top_p=1.0,
    max_tokens=500,
    stop=["# End", "\n\n\n"],  # Stop at code boundaries
    frequency_penalty=0.0,
    presence_penalty=0.0
)

Why these settings?

  • Temperature 0.1: Minimal randomness, maximum correctness
  • Stop sequences: Prevents generating extra unnecessary code
  • No penalties: Let the model use standard coding patterns

Scenario 5: Brainstorming Session

Goal: Maximum diversity of ideas


# Generate multiple diverse ideas by making multiple calls
ideas = []

for i in range(5):  # Get 5 different idea sets
    response = openai.ChatCompletion.create(
        model="gpt-4",
        messages=[{
            "role": "user",
            "content": "Suggest innovative mobile app ideas for fitness"
        }],
        temperature=1.0,       # High creativity
        top_p=0.95,           # Allow diverse tokens
        max_tokens=200,
        presence_penalty=0.7,  # High - want unique ideas each time!
        n=1
    )
    ideas.append(response.choices[0].message.content)

Why these settings?

  • Temperature 1.0: Good creative range
  • High presence penalty: Forces the model to explore different concepts each time
  • Multiple calls: Generates truly diverse idea sets

Advanced Topic: Sampling Methods Comparison 📊

Greedy Decoding vs Sampling

When temperature = 0 (or very close), models use greedy decoding: always pick the most probable next token.

Result: Completely deterministic. Same input = same output every time.

When temperature > 0, models use sampling: randomly pick from probability distribution.

Result: Different outputs each time (stochastic).

Which Sampling Method Should You Use?

Here's the decision tree for 2025:

🔵 Decision Guide:

Using Commercial API (OpenAI, Anthropic, Google)?

  • → Use Temperature + Top-P
  • → Set one to default (1.0), tune the other

Using Open-Source/Local Model?

  • → Use Temperature + Min-P (0.05-0.1)
  • → Avoid Top-P and Top-K

Need Absolute Determinism?

  • → Set Temperature = 0 (or use greedy decoding)
  • → Also consider setting a random seed for reproducibility

Common Mistakes to Avoid ⚠️

❌ Mistake 1: Tuning Temperature AND Top-P Together

Why it's bad: They interact unpredictably. You won't know which is causing behavior changes.

Fix: Pick one to tune. Set the other to 1.0.

❌ Mistake 2: Using High Temperature for Factual Tasks

Why it's bad: Introduces randomness where you need accuracy. Model might generate plausible-sounding but incorrect facts.

Fix: Use temperature ≤ 0.3 for any factual/technical content.

❌ Mistake 3: Setting Max Tokens Too Low

Why it's bad: Response gets cut off mid-sentence. Looks unprofessional and confusing.

Fix: Calculate expected response length + 20% buffer. Better to have room than to truncate!

❌ Mistake 4: Ignoring Context Window Limits

Why it's bad: Model silently truncates your prompt or errors out. You lose important context.

Fix: Always check: prompt_tokens + max_tokens ≤ context_window. Count tokens before sending!

❌ Mistake 5: Over-using Penalties

Why it's bad: Setting frequency/presence penalty too high (>1.0) makes text sound unnatural and forced.

Fix: Start small (0.1-0.3). Increase gradually only if needed. Rarely exceed 0.6.

❌ Mistake 6: Not Testing Different Settings

Why it's bad: You stick with defaults that might not be optimal for your specific use case.

Fix: Run A/B tests. Generate same prompt with different hyperparameters. Compare results.

Testing and Optimization Workflow 🔬

Here's a systematic approach to finding the best hyperparameters for your use case:

Step 1: Start with Baseline

Baseline Settings:
- Temperature: 0.7
- Top-P: 1.0 (or Min-P: 0.1 for open-source)
- Max tokens: Reasonable for your task
- Penalties: 0.0

Step 2: Identify Your Priority

Ask yourself: What matters most?

  • Accuracy/Correctness? → Lower temperature (0.1-0.3)
  • Creativity/Diversity? → Higher temperature (0.8-1.0)
  • Avoiding repetition? → Add frequency penalty (0.2-0.5)
  • Speed? → Reduce max_tokens
  • Cost? → Reduce max_tokens, optimize prompt length

Step 3: Adjust One Parameter at a Time

This is crucial! If you change multiple things, you won't know what helped.

Test 1: Baseline (temp=0.7)
Test 2: temp=0.3
Test 3: temp=0.5
Test 4: temp=0.9

Pick the best temperature, then move to next parameter.

Test 5: Best temp + frequency_penalty=0.2
Test 6: Best temp + frequency_penalty=0.4
...and so on

Step 4: Create a Test Suite


# Systematic hyperparameter testing
import openai

test_prompts = [
    "Explain quantum computing",
    "Write a creative story opening",
    "Debug this Python code: [code]"
]

temperatures = [0.2, 0.5, 0.7, 1.0]
results = {}

for temp in temperatures:
    results[temp] = []
    for prompt in test_prompts:
        response = openai.ChatCompletion.create(
            model="gpt-4",
            messages=[{"role": "user", "content": prompt}],
            temperature=temp,
            max_tokens=200
        )
        results[temp].append(response.choices[0].message.content)

# Now compare results[0.2] vs results[0.5] etc.
# Choose the temperature that gives best results across all prompts

Step 5: Measure Quality

Don't just eyeball results. Use metrics:

  • For factual accuracy: Compare against ground truth, count errors
  • For creativity: Measure diversity (unique words, sentence structures)
  • For user satisfaction: Run A/B tests with real users
  • For code: Run unit tests, measure syntax errors

Hyperparameter Cheat Sheet 📋

Parameter Range Use Low For Use High For
Temperature 0-2 Factual Q&A
Code generation
Technical docs
Translations
Creative writing
Brainstorming
Diverse outputs
Storytelling
Top-P 0-1 Precise answers
Focused output
Single best response
Exploring options
Creative tasks
Diverse vocabulary
Min-P 0.01-1 General purpose (0.1)
Focused output
Not recommended
(Use 0.05-0.1)
Top-K 1-∞ Very focused (5-10)
Deterministic
More variety (40-50)
(But use Top-P instead!)
Max Tokens 1-∞ Quick answers
Chat responses
Cost control
Long-form content
Articles
Detailed analysis
Frequency Penalty -2 to 2 Technical writing
(Need specific terms)
Avoiding repetition
Varied language
Creative content
Presence Penalty -2 to 2 Allow key term usage
Domain-specific text
Maximum diversity
Exploring topics
Idea generation

Quick Reference - Common Tasks 🎯

📝 Writing Blog Posts

temperature: 0.7
top_p: 1.0
max_tokens: 2000
frequency_penalty: 0.3
presence_penalty: 0.0

💻 Code Generation

temperature: 0.2
top_p: 1.0
max_tokens: 1000
frequency_penalty: 0.0
presence_penalty: 0.0

🎨 Creative Fiction

temperature: 0.9
top_p: 1.0
max_tokens: 3000
frequency_penalty: 0.3
presence_penalty: 0.1

❓ Q&A / Customer Support

temperature: 1.0
top_p: 0.7
max_tokens: 300
frequency_penalty: 0.0
presence_penalty: 0.0

📊 Data Analysis Reports

temperature: 0.3
top_p: 1.0
max_tokens: 1500
frequency_penalty: 0.1
presence_penalty: 0.0

💡 Brainstorming Ideas

temperature: 1.0
top_p: 0.95
max_tokens: 500
frequency_penalty: 0.0
presence_penalty: 0.6

Advanced Techniques for Experts 🚀

1. Adaptive Hyperparameter Selection

Instead of using fixed values, adjust hyperparameters based on the task:


def get_optimal_params(task_type):
    """Return hyperparameters optimized for task type"""
    
    configs = {
        "factual": {
            "temperature": 0.2,
            "top_p": 1.0,
            "max_tokens": 500,
            "frequency_penalty": 0.0
        },
        "creative": {
            "temperature": 0.9,
            "top_p": 1.0,
            "max_tokens": 2000,
            "frequency_penalty": 0.3
        },
        "code": {
            "temperature": 0.1,
            "top_p": 1.0,
            "max_tokens": 1500,
            "frequency_penalty": 0.0
        },
        "conversational": {
            "temperature": 1.0,
            "top_p": 0.8,
            "max_tokens": 300,
            "frequency_penalty": 0.0
        }
    }
    
    return configs.get(task_type, configs["conversational"])

# Usage
params = get_optimal_params("creative")
response = openai.ChatCompletion.create(
    model="gpt-4",
    messages=[{"role": "user", "content": "Write a story"}],
    **params  # Unpack the optimal parameters
)

2. Temperature Scheduling

For long-form generation, start high (creative) and gradually decrease (focused):


def generate_with_temperature_schedule(prompt, stages):
    """
    Generate text with decreasing temperature across stages
    stages: [(temperature, max_tokens), ...]
    """
    full_response = ""
    conversation = [{"role": "user", "content": prompt}]
    
    for temp, tokens in stages:
        response = openai.ChatCompletion.create(
            model="gpt-4",
            messages=conversation,
            temperature=temp,
            max_tokens=tokens
        )
        
        chunk = response.choices[0].message.content
        full_response += chunk
        
        # Add assistant's response to conversation
        conversation.append({"role": "assistant", "content": chunk})
        conversation.append({
            "role": "user", 
            "content": "Continue writing..."
        })
    
    return full_response

# Example: Start creative, end focused
result = generate_with_temperature_schedule(
    "Write a research paper about AI safety",
    stages=[
        (0.9, 500),  # Creative introduction
        (0.7, 1000), # Detailed body
        (0.3, 500)   # Precise conclusion
    ]
)

3. Multi-Sample Voting

Generate multiple responses with high temperature, then use a deterministic model to pick the best:


def generate_with_voting(prompt, n_samples=5):
    """Generate multiple samples and pick the best via voting"""
    
    # Step 1: Generate diverse candidates
    candidates = []
    for _ in range(n_samples):
        response = openai.ChatCompletion.create(
            model="gpt-4",
            messages=[{"role": "user", "content": prompt}],
            temperature=0.9,  # High diversity
            max_tokens=200
        )
        candidates.append(response.choices[0].message.content)
    
    # Step 2: Use low-temp model to pick best
    voting_prompt = f"""
    Here are {n_samples} candidate responses to: "{prompt}"
    
    Candidates:
    {chr(10).join(f"{i+1}. {c}" for i, c in enumerate(candidates))}
    
    Which response is the best? Reply with just the number (1-{n_samples}).
    """
    
    vote = openai.ChatCompletion.create(
        model="gpt-4",
        messages=[{"role": "user", "content": voting_prompt}],
        temperature=0.1,  # Deterministic voting
        max_tokens=10
    )
    
    best_idx = int(vote.choices[0].message.content.strip()) - 1
    return candidates[best_idx]

# Usage
best_response = generate_with_voting(
    "Explain quantum entanglement simply"
)

4. Context-Aware Max Tokens

Dynamically calculate max_tokens based on remaining context window:


import tiktoken  # OpenAI's tokenizer

def count_tokens(text, model="gpt-4"):
    """Count tokens in a text string"""
    encoding = tiktoken.encoding_for_model(model)
    return len(encoding.encode(text))

def smart_generate(messages, model="gpt-4", context_window=8192):
    """Automatically set max_tokens based on available context"""
    
    # Count tokens in messages
    prompt_tokens = sum(
        count_tokens(msg["content"], model) 
        for msg in messages
    )
    
    # Calculate safe max_tokens
    # Leave 10% buffer for safety
    available = context_window - prompt_tokens
    safe_max_tokens = int(available * 0.9)
    
    if safe_max_tokens < 100:
        raise ValueError("Not enough context window remaining!")
    
    response = openai.ChatCompletion.create(
        model=model,
        messages=messages,
        max_tokens=min(safe_max_tokens, 4000),  # Cap at 4000
        temperature=0.7
    )
    
    return response

# Usage - automatically handles token limits
response = smart_generate([
    {"role": "user", "content": "Your very long prompt here..."}
])

Troubleshooting Guide 🔧

Problem: Output is Too Random/Nonsensical

Symptoms: Text doesn't make sense, random words, incoherent sentences

Solutions:

  1. Reduce temperature from 1.0 → 0.7 → 0.5
  2. Reduce top_p from 0.95 → 0.8 → 0.6
  3. If using top-k, reduce from 50 → 30 → 10
  4. Check if you're accidentally using temperature > 1.5

Problem: Output is Too Repetitive/Boring

Symptoms: Same phrases over and over, monotonous language, always says the same thing

Solutions:

  1. Increase temperature from 0.5 → 0.7 → 0.9
  2. Add frequency_penalty: 0.2 → 0.4 → 0.6
  3. Try presence_penalty if frequency_penalty doesn't help
  4. Increase top_p to allow more token diversity

Problem: Responses Cut Off Mid-Sentence

Symptoms: Output ends abruptly, incomplete thoughts

Solutions:

  1. Increase max_tokens (double it as a test)
  2. Check you're not hitting context window limit
  3. Remove or adjust stop sequences if using them
  4. Calculate: prompt_tokens + max_tokens must fit in context_window

Problem: Output is Factually Incorrect

Symptoms: Plausible-sounding but wrong information, hallucinations

Solutions:

  1. Reduce temperature to ≤ 0.3 for factual tasks
  2. Use top_p around 0.1-0.3 instead of high values
  3. Add "Be factual and accurate" to system prompt
  4. Consider using retrieval-augmented generation (RAG) for better factual grounding

Problem: Responses Too Expensive

Symptoms: API costs are high

Solutions:

  1. Reduce max_tokens to minimum needed
  2. Optimize your prompts to be more concise
  3. Use a smaller model (e.g., GPT-3.5 instead of GPT-4) where appropriate
  4. Cache common responses
  5. Implement rate limiting

The Future of Hyperparameters (2025 and Beyond) 🔮

Emerging Trends

1. Min-P Becoming Standard

Expect commercial APIs (OpenAI, Anthropic) to add Min-P support soon. It's proven superior to Top-P in research.

2. Reasoning Models with Locked Parameters

Advanced reasoning models (like o1, o3) often lock their hyperparameters. You can't change temperature or top-p because they're optimized internally.

3. Automatic Hyperparameter Tuning

ML systems that learn optimal hyperparameters for your specific use case by analyzing your feedback.

4. Context-Aware Adaptive Sampling

Models that automatically adjust sampling strategy based on what they're generating (creative vs factual) without manual tuning.

Research to Watch

  • Top-n-sigma: Temperature-invariant sampling (published ACL 2025)
  • Mirostat: Adaptive sampling based on perplexity control
  • Dynamic Temperature: Automatically adjust temperature per-token
  • Beam Search Improvements: Better quality-diversity tradeoffs

Summary - Your Hyperparameter Journey 🎓

Congratulations! You've gone from zero knowledge to understanding hyperparameters like a pro!

Let's recap the essentials:

✅ Key Takeaways:
  • Temperature controls creativity (0.2 = factual, 0.9 = creative)
  • Top-P filters which words are considered (0.1 = focused, 0.95 = diverse)
  • Min-P is the new superior alternative to Top-P for open-source models
  • Max Tokens limits response length - plan it based on context window
  • Penalties fight repetition - use sparingly (0.1-0.5 range)
  • Tune ONE parameter at a time - never temperature AND top-p together
  • Different tasks need different settings - no one-size-fits-all
💡 Final Pro Tip:

Document your best configurations! Create a personal "cheat sheet" of what works for your specific use cases. This saves huge amounts of time later.

Further Learning Resources 📚

Want to dive deeper? Check out these resources:

Research Papers

  • "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs" (Nguyen et al., ICLR 2025)
  • "The Curious Case of Neural Text Degeneration" (Holtzman et al., 2020) - Introduced Top-P
  • "Top-n-sigma: Temperature-Invariant Sampling" (Tang et al., ACL 2025)

Documentation

  • OpenAI API Parameters: https://platform.openai.com/docs/api-reference/chat
  • Anthropic Claude Parameters: https://docs.anthropic.com/claude/reference
  • Hugging Face Generation: https://huggingface.co/docs/transformers/main_classes/text_generation

Conclusion 🎉

Hyperparameters are the secret sauce that transforms generic LLM outputs into exactly what you need.

Happy LLM Enginneering ! 🤖

Comments