Imagine you are baking a cake 🎂. You can adjust the oven temperature, the baking time, and the amount of sugar. Get these wrong — the cake burns, stays raw, or tastes terrible!
Hyperparameters in AI are exactly the same. They are the "recipe knobs" you set before training starts to get a great result.
What is Fine-Tuning?
Think of a very smart student who has read millions of books. They know a lot about the world — grammar, facts, reasoning. But they have never studied medical reports specifically.
Fine-tuning = taking that smart student and giving them a short, focused crash-course on medical reports so they become an expert in that one topic.
💡 Think of it like: A doctor who studied 10 years of general medicine, then does a 3-month specialisation in cardiology. They are still a doctor — but now much better at one specific thing!
Pre-Training vs Fine-Tuning
- Pre-Training = Huge dataset (internet text), done by Hugging Face / Google / Meta on thousands of GPUs over weeks
- Fine-Tuning = Your small dataset, same model, adjusted gently — done by YOU on 1 GPU in hours 🙌
- Result = A model that knows your domain and answers your specific questions!
How Training Actually Works
Before we look at hyperparameters, let's understand the training loop — what actually happens when a model learns. Think of it like learning to ride a bike 🚲:
- Step 1 — Feed Data: Give the model a batch of examples (sentences, questions, etc.)
- Step 2 — Make Predictions: Model guesses the answer (called the forward pass)
- Step 3 — Calculate Loss: How wrong was the guess? (the loss function measures this)
- Step 4 — Backpropagation: Figure out which parts of the model caused the mistake
- Step 5 — Update Weights: Nudge model parameters to reduce the mistake (this is where learning rate matters!)
- 🔁 Repeat: Do this for all batches, for multiple epochs
Every single hyperparameter controls one part of this loop. Now let's meet them one by one!
Hyperparameter 1 — Learning Rate
learning_rate = 1e-5
What is it?
The learning rate tells the model: "How big of a step should you take when you realise you made a mistake?"
💡 Analogy — Walking in the Dark:
- 🚶 Tiny step (low LR = 1e-5): Very slow, but you won't bang into walls
- 🏃 Giant leap (high LR = 0.1): Fast, but you might jump past the exit completely!
- 🎯 The sweet spot for fine-tuning Transformers is usually
1e-5to5e-5
What does 1e-5 actually mean?
1e-5 means 0.00001. That is a tiny, tiny nudge.
Researchers found through years of experiments that pre-trained Transformer weights are very delicate. Big updates would make the model "forget" everything it learned during pre-training. This is called catastrophic forgetting.
So we use a very small learning rate to gently adjust the weights without destroying what the model already knows.
What happens with different learning rates?
1e-2(too high) → Model overshoots, loss bounces wildly, never converges 💥1e-4(high-ish) → Learns fast but can overshoot, risky for fine-tuning ⚠️1e-5(sweet spot ✅) → Steady learning, model improves smoothly1e-7(too low) → Model barely moves, training takes forever ⏳
Can you change it?
✅ YES — you should experiment! Common values for fine-tuning: 2e-5, 3e-5, 5e-5. Start with 2e-5 and adjust based on results.
Tip — Learning Rate Schedulers
Nobody uses a fixed learning rate throughout training. We use a scheduler that changes the learning rate automatically:
- Cosine with warmup — Most popular. Slowly goes up, then gently comes down like a hill 🏔️
- Linear decay — Starts at your LR, decreases to zero in a straight line
- WSD (Warmup-Stable-Decay) — New in 2024–2026, used by Mistral and Llama teams
Hyperparameter 2 — Batch Size
batch_size = 32
What is it?
Imagine you are a teacher grading exams. You could grade 1 paper at a time and update your rubric after each one — very slow! Or you could grade 32 papers together, look at patterns, and update your rubric once. That is the batch size!
💡 Analogy — Grocery Shopping 🛒:
- Batch size = 1: You go to the store, buy 1 item, come home, then go again. Very slow!
- Batch size = 32: You buy 32 items in one trip. Efficient!
- Batch size = 1024: You need a truck! (requires a lot of GPU memory)
How is 32 decided?
Batch size of 32 is a classic default that fits comfortably in most GPUs. The right choice depends on:
- 🖥️ Your GPU memory — More memory = bigger batch possible
- 📊 Your dataset size — Tiny dataset? Use smaller batches
- 🎯 Task type — Some tasks work better with smaller, noisier gradients
Quick reference by GPU size
- 8 GB GPU (e.g. T4) → batch size 8–16
- 16 GB GPU (e.g. V100) → batch size 16–32
- 40 GB GPU (A100) → batch size 32–64
- 80 GB GPU (A100 large) → batch size 64–128
The Secret Trick — Gradient Accumulation 🪄
What if your GPU can only fit batch size 8, but you want the effect of batch size 32?
Use gradient accumulation! Run 4 small batches of 8, add up (accumulate) the gradients, and then update the model — as if you used batch size 32. No extra GPU memory needed! 🎉
gradient_accumulation_steps = 4 # batch 8 × 4 steps = effective batch size 32
Can you change it?
✅ YES. Start with what fits your GPU. Use gradient accumulation to simulate larger batches when needed.
Hyperparameter 3 — Warmup Steps
warmup = 600
What is it?
Warmup steps tell the model: "Start with a very tiny learning rate, and slowly increase it to the full learning rate over the first 600 steps."
💡 Analogy — Car Engine in Winter ❄️:
You don't slam the accelerator on a cold engine — it stalls! You warm it up slowly first, then go full speed. Warmup steps do the same for your model — they prevent huge, damaging weight updates right at the start of training.
How is 600 calculated?
A common rule of thumb: warmup_steps = total_training_steps × 0.06
Here is an example calculation:
Dataset: 10,000 examples
Batch size: 32
Steps per epoch: 10,000 ÷ 32 = 312 steps
3 Epochs: 312 × 3 = 936 total steps
Warmup (6%): 936 × 0.06 ≈ 56 steps
(600 is used as a fixed default when total steps are larger, e.g. 10,000+ steps)
What happens step by step during warmup?
- Step 1 → LR = 0.000001 (very tiny, just starting)
- Step 100 → LR = 0.000003 (slowly climbing)
- Step 300 → LR = 0.000007 (almost there)
- Step 600 → LR = 0.00001 (full learning rate reached! 🎯)
- After step 600 → LR gently decays (cosine or linear)
Can you change it?
✅ YES. General rule: warmup = 5–10% of total training steps. Or simply use warmup_ratio=0.06 in Hugging Face and it calculates warmup steps automatically!
Hyperparameter 4 — Max Sequence Length
max_seq_length = 128
What is it?
Transformers don't read sentences like humans — they read tokens (chunks of words). max_seq_length tells the model: "Never read more than 128 tokens at a time."
If a sentence is longer → it gets cut off (truncated). If shorter → it gets padded with special empty tokens to fill up to 128.
💡 Analogy — Reading a Book 📖:
- Imagine you can only read 128 words per page at a time
- If the chapter has 500 words → it gets cut after 128 words
- If the chapter has 50 words → you add blank lines to fill the page (padding)
What is a token exactly?
A token is roughly ¾ of a word on average. So max_seq_length = 128 handles about 90–100 words. Here are some examples:
"Hello world" → 2 tokens
"I love pizza" → 3 tokens
"unbelievable" → 3 tokens (un + belie + vable)
"GPT-4" → 3 tokens (G + PT + -4)
Common values by task
- Sentiment / Tweet classification: 64–128 tokens (short inputs)
- Q&A / NLI / Classification: 128–256 tokens (medium inputs)
- Document summarisation: 512–1024 tokens (long documents)
- Modern LLMs like Llama 3 / Gemma : 8,192–131,072 tokens (very long context)
⚠️ Important Warning — Memory grows quadratically!
Doubling max_seq_length from 128 → 256 uses 4× more memory. Going from 128 → 512 uses 16× more memory!
Always start small and only increase if your task actually needs it.
Can you change it?
✅ YES. Run this on a few samples to see how long your inputs actually are, then set max_seq_length just above the 95th percentile of that:
lengths = [len(tokenizer(text)['input_ids']) for text in your_texts]
print(f"Max length: {max(lengths)}")
print(f"95th percentile: {sorted(lengths)[int(len(lengths)*0.95)]}")
Hyperparameter 5 — Number of Training Epochs
num_train_epochs = 3.0
What is it?
An epoch means the model sees every single training example once. num_train_epochs = 3 means the model goes through the entire dataset 3 full times.
💡 Analogy — Studying for an Exam 📚:
- Epoch 1: You read the textbook once — you understand the basics
- Epoch 2: You read it again — more things click, you remember better
- Epoch 3: One more read — you feel confident and ready for the exam ✅
- Epoch 10: You've read it so many times you're just memorising word for word (overfitting!) — not great ❌
What happens with too few or too many epochs?
- 1–2 epochs: Underfit — model hasn't learned enough, poor accuracy
- 3–5 epochs ✅: Sweet spot — good generalisation, works well on new unseen data
- 10+ epochs: Overfit — memorises training data, fails badly on new data
Can you change it?
✅ YES — but use Early Stopping! Instead of guessing the right epoch count, set load_best_model_at_end=True together with an evaluation strategy. Training automatically keeps the best checkpoint. This is the best practice — no wasted compute!
Other Important Hyperparameters (Full List)
The five values above are just the beginning. Here are the other hyperparameters you will regularly encounter when working with Hugging Face Transformers:
Regularisation — Preventing Overfitting
weight_decay = 0.01— Gently pushes model weights toward zero during training. Prevents any single weight from becoming too dominant and memorising training quirks. Think of it as a "tax on complexity" ⚖️dropout = 0.1— Randomly switches off 10% of neurons each training step. Forces the model to learn multiple pathways to the right answer, not rely on just one neuron
💡 Dropout Analogy — Team Training 🏀: A basketball coach randomly benches 1–2 players each practice. This forces the whole team to learn every position, not rely on just one star player. Dropout does the same for neurons!
Speed and Memory
fp16 = True— Use 16-bit floating point numbers instead of 32-bit. Makes training roughly 2× faster and uses half the memory. Great for T4 and V100 GPUsbf16 = True— A more numerically stable version of fp16. The default for A100 and H100 GPUs. Prefer this over fp16 when your GPU supports itgradient_checkpointing = True— Saves GPU memory by recomputing some calculations during backpropagation instead of storing them. Slightly slower but enables training on much smaller GPUsgradient_accumulation_steps— Simulates a larger batch size without needing more memory (covered above!)
Saving and Evaluation
eval_steps = 500— Evaluate the model on your validation set every 500 training stepssave_steps = 500— Save a model checkpoint every 500 steps so you never lose progressload_best_model_at_end = True— At the end of training, automatically restore the best checkpoint rather than just the last onemetric_for_best_model = "f1"— Which metric to use when deciding what is "best" (accuracy, f1, loss, etc.)
Reproducibility
seed = 42— Sets a random seed so you get the same results every time you run the same code. The entire ML world loves the number 42 for this! 😄
LoRA - Fine-Tuning Revolution ⚡
Traditionally, fine-tuning updates all the billions of parameters in a model. That requires massive GPUs and huge memory — not practical for most people!
In 2021, researchers invented LoRA (Low-Rank Adaptation) — a way to fine-tune with 100× fewer trainable parameters by only training small "adapter" matrices that sit next to the original frozen weights.
💡 Analogy — Home Renovation 🏠: Instead of rebuilding the entire house, you only renovate the kitchen. Same house — but now much better at cooking!
How LoRA works
- Original Weights: Frozen — not updated at all during training
- LoRA Adapters: Tiny matrices added beside the original weights — only these are trained
- Result: Same fine-tuning quality, a tiny fraction of the training cost!
LoRA Hyperparameters
lora_r = 8— Rank of the adapter matrix. Higher = more capacity but more memory. Common values: 4, 8, 16, 64lora_alpha = 16— Scaling factor for LoRA updates. Usually set to 2× the rank valuelora_dropout = 0.05— Slight dropout applied inside the adapters to prevent overfittingtarget_modules— Which layers to apply LoRA to. Usually the attention query and value layers
✅ LoRA (and upgraded versions QLoRA and DoRA) are the standard for fine-tuning LLMs on consumer hardware. You can fine-tune a 7 billion parameter model on a single 16 GB GPU using QLoRA!
Full Code Example — Step by Step
Let's put everything together with a real working fine-tuning script. We will fine-tune BERT for sentiment analysis (Positive / Negative movie reviews) using the IMDB dataset.
Step 1: Install the Libraries
📌 What this code does: Installs all the Python packages we need — like buying all your cooking ingredients before you start cooking! This only needs to be run once.
pip install transformers datasets peft trl accelerate bitsandbytes -q
Step 2: Load the Model and Tokenizer
📌 What this code does: Downloads a pre-trained BERT model from Hugging Face (like downloading a very smart brain 🧠). Also loads the tokenizer — the tool that converts your words into numbers the model understands. Think of the tokenizer as a translator between human language and model language.
from transformers import AutoModelForSequenceClassification, AutoTokenizer
# The model we will fine-tune
model_name = "bert-base-uncased"
# Load the tokenizer — converts words into numbers
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Load the model with 2 output labels: Positive / Negative
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
num_labels=2
)
print(f"Model loaded! Parameters: {sum(p.numel() for p in model.parameters()):,}")
Output:
Model loaded! Parameters: 109,483,778
That is 109 million parameters — the brain we are about to fine-tune! 🧠
Step 3: Load and Tokenize the Dataset
📌 What this code does: Loads 50,000 IMDB movie reviews (already labelled Positive or Negative), then converts all the text into tokens that the model can read. The max_length=128 here is our max_seq_length hyperparameter — we only look at the first 128 tokens of each review.
from datasets import load_dataset
# Load IMDB dataset — 25,000 train + 25,000 test reviews
dataset = load_dataset("imdb")
def tokenize_function(examples):
return tokenizer(
examples["text"],
max_length=128, # ← our max_seq_length hyperparameter!
truncation=True, # cut off anything beyond 128 tokens
padding="max_length" # pad shorter reviews to reach exactly 128
)
# Apply tokenization to the entire dataset at once
tokenized_dataset = dataset.map(tokenize_function, batched=True)
print("Tokenization complete! ✅")
Step 4: Set All the Hyperparameters
📌 What this code does: This is the most important step! We set all our hyperparameters in one place using TrainingArguments. Every single value here directly controls how the model learns. Think of this as setting all the dials on your oven before putting the cake in 🎂 — temperature, timer, fan speed, all of it!
from transformers import TrainingArguments
training_args = TrainingArguments(
# Where to save the model
output_dir="./results",
# How many times to go through the full dataset
num_train_epochs=3, # ← our num_train_epochs hyperparameter!
# How many samples per GPU per step
per_device_train_batch_size=32, # ← our batch_size hyperparameter!
per_device_eval_batch_size=32,
# Speed of weight updates
learning_rate=1e-5, # ← our learning_rate hyperparameter!
# Steps to slowly ramp up the learning rate from zero
warmup_steps=600, # ← our warmup hyperparameter!
# How the LR changes over time after warmup ends
lr_scheduler_type="cosine",
# Prevents overfitting — a light tax on large weights
weight_decay=0.01,
# How often to check performance on the validation set
evaluation_strategy="steps",
eval_steps=500,
# How often to save a model checkpoint
save_steps=500,
# At the end, restore the best checkpoint automatically
load_best_model_at_end=True,
# Use 16-bit floats for faster training
fp16=True,
# Makes results reproducible every run
seed=42,
# Where to save training logs
logging_dir="./logs",
logging_steps=100,
)
print("Training arguments configured! ✅")
Step 5: Train the Model
📌 What this code does: Creates the Hugging Face Trainer — an all-in-one training engine. You give it the model, the hyperparameters, and the data. It handles everything else automatically: the training loop, evaluation at each checkpoint, and saving the best model. Think of it as a very capable robot chef who follows your recipe perfectly! 🤖
from transformers import Trainer
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
def compute_metrics(eval_pred):
# This runs after each evaluation step and returns accuracy + F1 score
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1) # pick the class with the highest score
accuracy = accuracy_score(labels, predictions)
f1 = f1_score(labels, predictions, average="weighted")
return {"accuracy": accuracy, "f1": f1}
# Assemble the Trainer with all our ingredients
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized_dataset["train"],
eval_dataset=tokenized_dataset["test"],
compute_metrics=compute_metrics,
)
# Start training! 🚀
trainer.train()
# Evaluate on the test set when done
results = trainer.evaluate()
print(f"Final Test Accuracy: {results['eval_accuracy']:.4f}")
print(f"Final F1 Score: {results['eval_f1']:.4f}")
Output:
Training complete! 🎉
Final Test Accuracy: 0.9312
Final F1 Score: 0.9311
93% accuracy — and we only fine-tuned for 3 epochs! 🌟
Step 6 (Bonus) — Fine-Tune with LoRA on a Small GPU
📌 What this code does: Instead of training all 109 million parameters (expensive and memory-hungry!), we add tiny LoRA adapters and only train those — roughly 300,000 parameters. This makes fine-tuning possible on a small, cheap GPU. Think of it as renovating only the kitchen of a house instead of rebuilding the entire building 🏠
from peft import LoraConfig, get_peft_model, TaskType
# Define LoRA configuration
lora_config = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=8, # ← lora_r: rank of the adapter
lora_alpha=16, # ← lora_alpha: scaling factor (2 × rank)
lora_dropout=0.05, # ← slight dropout inside adapters
target_modules=["query", "value"], # apply LoRA only to attention layers
bias="none"
)
# Wrap the original model with LoRA adapters
lora_model = get_peft_model(model, lora_config)
# See exactly how few parameters we now train!
lora_model.print_trainable_parameters()
Output:
trainable params: 296,448 || all params: 109,780,226 || trainable%: 0.27%
We are training only 0.27% of the total parameters — but we get almost the same accuracy! 🎉
# Use the exact same TrainingArguments and Trainer — works identically!
lora_trainer = Trainer(
model=lora_model,
args=training_args,
train_dataset=tokenized_dataset["train"],
eval_dataset=tokenized_dataset["test"],
compute_metrics=compute_metrics,
)
lora_trainer.train()
print("LoRA fine-tuning complete! 🎉")
How to Choose the Right Values
Choosing hyperparameters is part science, part art. Here is a practical decision guide you can follow for almost any project:
How much GPU memory do you have?
- Less than 16 GB → Use LoRA + gradient accumulation + batch size 8–16
- 16–40 GB → Standard fine-tuning, batch size 16–32
- 40 GB or more → Increase batch size to 64+, can train the full model
How big is your dataset?
- Less than 1,000 examples → 5 epochs, small LR (1e-5), add more dropout
- 1,000 – 50,000 examples → 3 epochs, LR 2e-5, standard settings
- 50,000+ examples → 1–3 epochs, LR 3e-5, can use larger batches
How long is your input text?
- Short (tweets, labels) →
max_seq_length64–128 - Medium (paragraphs, Q&A) →
max_seq_length256–512 - Long (documents, reports) →
max_seq_length1024–4096
Is the model overfitting or underfitting?
- Overfitting (train loss low, validation loss high) → Add more dropout, reduce epochs, increase weight decay
- Underfitting (both losses still high) → Add more epochs, increase LR slightly, consider a bigger model
Common Mistakes — What NOT to Do 🚫
🚫 Don't use a high learning rate for fine-tuning. Using 0.1 or even 0.001 with Transformers causes catastrophic forgetting — the model loses all its pre-training knowledge instantly.
🚫 Don't train for too many epochs on a small dataset. If you have 500 examples and train for 20 epochs, the model memorises every single example and fails badly on new data it hasn't seen before.
🚫 Don't skip warmup steps. Without warmup, the learning rate hits full speed at step 1 — like slamming the accelerator on a cold engine. Early gradients can be huge and unstable.
🚫 Don't set max_seq_length larger than you need. Memory grows quadratically with sequence length. Using 1024 when 128 is enough wastes 64× the memory and makes training painfully slow.
🚫 Don't only watch training loss. Always monitor your validation loss too. If training loss keeps going down but validation loss goes up — you are overfitting. Stop early!
Good Practices — What TO Do ✅
✅ Always use a learning rate scheduler. Set lr_scheduler_type="cosine" and warmup_ratio=0.06 as your starting point for almost any fine-tuning job.
✅ Use bf16=True on A100/H100 GPUs. It is more numerically stable than fp16 and is the default for large model training.
✅ Use gradient accumulation on small GPUs. gradient_accumulation_steps=4 with batch_size=8 gives you an effective batch of 32 with no extra memory needed.
✅ Always set a random seed. seed=42 makes your results reproducible — you and your team get identical numbers every single run.
✅ Track your experiments. Use Weights & Biases (report_to="wandb") to log every run.This is how professional ML engineers compare hyperparameter combinations without losing track of what they tried.
Your Starter Config — Copy and Paste This 📋
📌 What this code does: A ready-to-use, sensible starting configuration for almost any text classification fine-tuning task. Copy it, paste it, then adjust the values based on your GPU, dataset size, and task. Change one thing at a time!
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./my_model",
# Core training
num_train_epochs=3,
per_device_train_batch_size=16, # adjust based on your GPU
per_device_eval_batch_size=16,
gradient_accumulation_steps=2, # effective batch size = 16 x 2 = 32
# Learning rate
learning_rate=2e-5,
lr_scheduler_type="cosine",
warmup_ratio=0.06, # auto-calculates warmup steps as 6% of total
# Regularisation
weight_decay=0.01,
# Evaluation and saving
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1",
# Speed and memory
bf16=True, # use fp16=True if no A100/H100
gradient_checkpointing=True,
# Reproducibility
seed=42,
)
And your tokenizer call — always match max_length to your task:
tokenizer(text, max_length=128, truncation=True, padding="max_length")
Quick Summary 📝
What we learned :
learning_rate = 1e-5→ How big each learning step is. Too high = overshoots and forgets. Too low = too slow. Sweet spot for Transformers: 1e-5 to 5e-5batch_size = 32→ How many examples processed at once. Bigger = faster but needs more GPU memory. Use gradient accumulation when memory is tightwarmup = 600→ Steps to slowly ramp the LR up from zero. Prevents unstable early training. Rule of thumb: ~6% of total training stepsmax_seq_length = 128→ Max tokens read per input. Memory grows quadratically — always start small and only increase when needednum_train_epochs = 3.0→ Full passes through your data. 3–5 is the sweet spot. More risks overfitting, fewer risks underfitting
Happy fine-tuning! 🤗✨
Comments
Post a Comment