Imagine you want to become a chef. You have two choices:
Option A: Start from scratch. Learn what fire is. Learn how to hold a knife. Learn every spice from zero. Takes years!
Option B: Find a world-class chef who already knows everything. Ask them to teach you just the last part — your specific cuisine. Done in weeks!
Instead of training a neural network from zero (which takes weeks and massive compute), you borrow a model that was already trained by Google, Meta, or OpenAI on millions of images or texts. Then you either use it directly or teach it your specific task. This is called Transfer Learning — and it's the #1 superpower of modern Deep Learning. ⚡
🧠 Section 1: What is Transfer Learning? (The Foundation)
Every time a neural network trains on data, it doesn't just memorize answers. It builds internal knowledge — patterns, shapes, textures, relationships — stored inside its millions of weight parameters.
Transfer Learning is the idea of taking that stored knowledge and moving it to a new problem. You're not starting from zero. You're standing on the shoulders of giants. 🏔️
🧩 A Simple Analogy: The Expert Photographer
Say you're a professional wildlife photographer for 10 years. You've mastered lighting, composition, timing, and camera settings.
Now your friend asks you to photograph their wedding. You don't need to re-learn what a camera is! You transfer all your existing skills and just learn the specific style of wedding photography. Done fast. Done well.
Neural networks do the same thing. A model trained on 14 million images has learned to recognize edges, colors, textures, faces, objects. These are universal visual skills. You can borrow them for your specific task — whether it's detecting skin diseases or counting cars.
🏗️ Section 2: What Are Pretrained Models? How Were They Built?
A pretrained model is a neural network that has already been trained on a massive dataset — usually by a research lab or big tech company — and whose learned weights are made available to the public.
📦 What's Inside a Pretrained Model?
Think of a deep neural network like a multi-story building. Each floor (layer) is responsible for learning something different:
→ [Floor 3: Textures & Patterns] → [Floor 4: Parts (eyes, wheels, leaves)]
→ [Floor 5: Objects (face, car, tree)] → [Roof: Final Classification]
The lower floors learn general, universal features that are useful for any image task. The higher floors learn task-specific features that only matter for the original training task (e.g., identifying ImageNet classes).
When you borrow a pretrained model, you get ALL these floors for free! You only need to decide what to do with the roof.
🌍 Where Do These Models Come From?
- ImageNet Challenge (ILSVRC): Models trained on 1.2 million images across 1,000 categories. Gave us ResNet, VGG, Inception, EfficientNet.
- Google's JFT-300M / JFT-3B: 3 billion images! Used to train Vision Transformers (ViT) — the current state of the art.
- Common Crawl (for NLP): Petabytes of web text used to train BERT, GPT, LLaMA, Mistral and other language models.
- LAION-5B (for Multimodal): 5 billion image-text pairs used to train CLIP, DALL-E, Stable Diffusion.
Training GPT-4 from scratch is estimated to cost over $100 million in compute. Fine-tuning it for your task might cost $50–$500 on cloud GPUs. Transfer learning is not just convenient — it's economically transformative! 💰
⚖️ Section 3: The Two Big Strategies
Once you have a pretrained model, you have two main ways to use it. Understanding the difference between them is the most important thing in this guide.
Strategy 1: Feature Extraction 🧲
You use the pretrained model as a fixed feature extractor. You freeze ALL its weights — nothing changes inside the pretrained model. You only add and train a small new "head" (classifier) on top.
│ PRETRAINED MODEL (ALL FROZEN ❄️) │
│ Conv1 → Conv2 → Conv3 → ... → Pool │
│ (weights do NOT change at all) │
└─────────────────────────────────────────┘
↓
Feature Vector (e.g., 2048 numbers)
↓
┌─────────────────────────────────────────┐
│ YOUR NEW HEAD (Trainable 🔥) │
│ Dense(256) → Dropout → Dense(N classes)│
└─────────────────────────────────────────┘
Think of it like: You hire a brilliant analyst (the pretrained model) who generates a detailed report (feature vector) on every piece of data. You then train a simple decision-maker on those reports. The analyst never changes. You only train the decision-maker.
Strategy 2: Fine-Tuning 🎛️
You unfreeze some or all layers of the pretrained model and let them update their weights on your new data — but with a very small learning rate so you don't destroy the existing knowledge.
│ PRETRAINED MODEL │
│ Conv1 → Conv2 (FROZEN ❄️) │
│ Conv3 → Conv4 → Conv5 (TRAINABLE 🔥) │
└─────────────────────────────────────────┘
↓
┌─────────────────────────────────────────┐
│ YOUR NEW HEAD (Trainable 🔥) │
│ Dense(256) → Dropout → Dense(N classes)│
└─────────────────────────────────────────┘
Think of it like: Now you let the expert analyst actually adjust their analysis style to better fit your specific domain. They don't start from scratch — they gently adapt their deep expertise to your new context.
📊 Feature Extraction vs Fine-Tuning: Side-by-Side
| Aspect | Feature Extraction | Fine-Tuning |
|---|---|---|
| Pretrained weights | Completely frozen ❄️ | Partially or fully unfrozen 🔥 |
| Training speed | Very fast ⚡ | Slower (more params) |
| Data needed | Works with very little (100–1000 samples) | Needs more (1000–10000+ samples) |
| Risk of overfitting | Lower | Higher (if data is small) |
| Final accuracy | Good | Usually better 🏆 |
| GPU memory | Lower | Higher |
| Best when... | Small dataset, similar domain | Larger dataset, different domain |
🧲 Section 4: Feature Extraction — Deep Dive with Code
Let's build a real image classifier using Feature Extraction with ResNet50 (a classic, battle-tested pretrained model).
Step 1: Load the Pretrained Model
import torch
import torch.nn as nn
import torchvision.models as models
import torchvision.transforms as transforms
from torch.utils.data import DataLoader
from torchvision.datasets import ImageFolder
# Load ResNet50 pretrained on ImageNet
# weights='IMAGENET1K_V1' loads the classic pretrained weights
model = models.resnet50(weights='IMAGENET1K_V1')
print("ResNet50 loaded!")
print(f"Total parameters: {sum(p.numel() for p in model.parameters()):,}")
Output:
ResNet50 loaded!
Total parameters: 25,557,032
Step 2: Freeze ALL Pretrained Layers
# Freeze every single layer in the pretrained model.
# requires_grad=False means: "Don't update these weights during training"
for param in model.parameters():
param.requires_grad = False
# Verify: count trainable parameters (should be 0 now)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Trainable parameters after freezing: {trainable:,}")
# Output: Trainable parameters after freezing: 0
Step 3: Replace the Final Classification Head
# ResNet50's final layer is model.fc
# By default it outputs 1000 classes (ImageNet).
# We replace it with our own classifier for OUR number of classes.
NUM_CLASSES = 10 # Example: 10 dog breeds
# The original model.fc takes 2048 features as input
in_features = model.fc.in_features # = 2048
# Build a custom classification head
# This head IS trainable — only these params will update!
model.fc = nn.Sequential(
nn.Linear(in_features, 256),
nn.ReLU(),
nn.Dropout(p=0.4),
nn.Linear(256, NUM_CLASSES)
)
# Verify: now only the new head is trainable
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"Trainable params: {trainable:,}")
print(f"Total params: {total:,}")
print(f"Training only {trainable/total*100:.2f}% of the model!")
Output:
Trainable params: 524,810
Total parameters: 26,081,842
Training only 2.01% of the model!
You're training just 2% of the total parameters. Training is lightning fast and almost impossible to overfit, even on very small datasets. 🏎️
Step 4: Prepare Data with Correct Transforms
# CRITICAL: ImageNet pretrained models expect specific input preprocessing.
# Always normalize with ImageNet mean and std — the model was trained with these!
IMAGENET_MEAN = [0.485, 0.456, 0.406]
IMAGENET_STD = [0.229, 0.224, 0.225]
train_transforms = transforms.Compose([
transforms.Resize((256, 256)),
transforms.RandomCrop(224), # ResNet expects 224x224
transforms.RandomHorizontalFlip(), # Basic augmentation
transforms.ColorJitter(brightness=0.2, contrast=0.2),
transforms.ToTensor(),
transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD)
])
val_transforms = transforms.Compose([
transforms.Resize((224, 224)), # No augmentation for validation
transforms.ToTensor(),
transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD)
])
# Load your dataset (folder structure: data/train/class1, data/train/class2, ...)
train_dataset = ImageFolder('data/train', transform=train_transforms)
val_dataset = ImageFolder('data/val', transform=val_transforms)
train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True, num_workers=4)
val_loader = DataLoader(val_dataset, batch_size=32, shuffle=False, num_workers=4)
print(f"Training samples: {len(train_dataset)}")
print(f"Validation samples: {len(val_dataset)}")
print(f"Classes: {train_dataset.classes}")
Step 5: Train ONLY the New Head
import torch.optim as optim
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = model.to(device)
# Only pass trainable parameters to the optimizer
# This is important! Only model.fc params are trainable.
optimizer = optim.Adam(
filter(lambda p: p.requires_grad, model.parameters()),
lr=1e-3
)
criterion = nn.CrossEntropyLoss()
scheduler = optim.lr_scheduler.StepLR(optimizer, step_size=5, gamma=0.5)
def train_one_epoch(model, loader, optimizer, criterion, device):
model.train()
total_loss, correct, total = 0, 0, 0
for images, labels in loader:
images, labels = images.to(device), labels.to(device)
optimizer.zero_grad()
outputs = model(images)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
total_loss += loss.item()
preds = outputs.argmax(dim=1)
correct += (preds == labels).sum().item()
total += labels.size(0)
return total_loss / len(loader), correct / total * 100
def evaluate(model, loader, criterion, device):
model.eval()
total_loss, correct, total = 0, 0, 0
with torch.no_grad():
for images, labels in loader:
images, labels = images.to(device), labels.to(device)
outputs = model(images)
loss = criterion(outputs, labels)
total_loss += loss.item()
preds = outputs.argmax(dim=1)
correct += (preds == labels).sum().item()
total += labels.size(0)
return total_loss / len(loader), correct / total * 100
# Training loop
EPOCHS = 10
best_val_acc = 0.0
print(f"Training on: {device}")
print("="*55)
for epoch in range(1, EPOCHS + 1):
train_loss, train_acc = train_one_epoch(
model, train_loader, optimizer, criterion, device
)
val_loss, val_acc = evaluate(model, val_loader, criterion, device)
scheduler.step()
print(f"Epoch {epoch:02d}/{EPOCHS} | "
f"Train Loss: {train_loss:.4f} | Train Acc: {train_acc:.2f}% | "
f"Val Loss: {val_loss:.4f} | Val Acc: {val_acc:.2f}%")
# Save the best model
if val_acc > best_val_acc:
best_val_acc = val_acc
torch.save(model.state_dict(), 'best_model_feature_extraction.pth')
print(f" ✅ New best model saved! (Val Acc: {val_acc:.2f}%)")
print(f"\n🏆 Best Validation Accuracy: {best_val_acc:.2f}%")
🎛️ Section 5: Fine-Tuning — Deep Dive with Code
Fine-tuning takes feature extraction one step further. After the new head has warmed up, you gently unfreeze deeper layers and let the whole model adapt to your data.
Phase 1 (Feature Extraction): Train ONLY the new head for a few epochs. This "warms up" the head so it produces reasonable outputs.
Phase 2 (Fine-Tuning): Unfreeze some or all of the pretrained layers. Continue training with a MUCH smaller learning rate. Now the whole network adapts together.
Why two phases? If you fine-tune from the very start, the large gradients from the random new head can destroy the pretrained weights. Warm up first — then unfreeze!
Phase 1: Warm Up the Head (Same as Feature Extraction)
import torch
import torch.nn as nn
import torchvision.models as models
import torch.optim as optim
# Load pretrained EfficientNet-B0 (excellent for efficiency)
model = models.efficientnet_b0(weights='IMAGENET1K_V1')
# --- PHASE 1: Freeze everything ---
for param in model.parameters():
param.requires_grad = False
# Replace classifier head
# EfficientNet's classifier is model.classifier[1]
in_features = model.classifier[1].in_features # = 1280
NUM_CLASSES = 5
model.classifier = nn.Sequential(
nn.Dropout(p=0.3, inplace=True),
nn.Linear(in_features, 256),
nn.ReLU(),
nn.Dropout(p=0.3),
nn.Linear(256, NUM_CLASSES)
)
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = model.to(device)
# Phase 1 optimizer — high learning rate is OK since only head trains
optimizer_phase1 = optim.Adam(
filter(lambda p: p.requires_grad, model.parameters()),
lr=1e-3
)
criterion = nn.CrossEntropyLoss()
print("Phase 1: Training new head only...")
# (Run train_one_epoch for 5 epochs here — same function as above)
# ... (training loop same as feature extraction section)
print("Phase 1 complete! Head is warmed up ✅")
Phase 2: Unfreeze & Fine-Tune
# --- PHASE 2: Selectively unfreeze layers ---
# Strategy: unfreeze ONLY the last few blocks of EfficientNet
# This gives us the benefits of fine-tuning without catastrophic forgetting
# First unfreeze everything
for param in model.parameters():
param.requires_grad = True
# Then re-freeze the early layers (feature 0 through 5 = early blocks)
# We keep features 6, 7, 8 (deeper blocks) + classifier trainable
layers_to_freeze = list(model.features.children())[:6] # First 6 blocks
for layer in layers_to_freeze:
for param in layer.parameters():
param.requires_grad = False
# Count trainable parameters now
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"Phase 2: Training {trainable:,} / {total:,} parameters")
print(f" ({trainable/total*100:.1f}% of model)")
# CRITICAL: Use a much smaller learning rate for fine-tuning!
# Large LR will destroy pretrained knowledge. Rule of thumb: 10x–100x smaller.
optimizer_phase2 = optim.Adam([
# Deeper pretrained layers — very small LR
{'params': model.features[6:].parameters(), 'lr': 1e-5},
# New classifier head — slightly larger LR
{'params': model.classifier.parameters(), 'lr': 1e-4},
])
# Cosine annealing LR scheduler — standard in fine-tuning
scheduler_phase2 = optim.lr_scheduler.CosineAnnealingLR(
optimizer_phase2, T_max=15, eta_min=1e-7
)
print("\nPhase 2: Fine-tuning with differential learning rates...")
EPOCHS_PHASE2 = 15
best_val_acc = 0.0
for epoch in range(1, EPOCHS_PHASE2 + 1):
train_loss, train_acc = train_one_epoch(
model, train_loader, optimizer_phase2, criterion, device
)
val_loss, val_acc = evaluate(model, val_loader, criterion, device)
scheduler_phase2.step()
print(f"Epoch {epoch:02d}/{EPOCHS_PHASE2} | "
f"Train: {train_acc:.2f}% | Val: {val_acc:.2f}%")
if val_acc > best_val_acc:
best_val_acc = val_acc
torch.save(model.state_dict(), 'best_model_finetuned.pth')
print(f" ✅ New best: {val_acc:.2f}%")
print(f"\n🏆 Fine-tuned Best Val Accuracy: {best_val_acc:.2f}%")
❄️ Section 6: Layer Freezing — What, Why, and How
Layer freezing is the core mechanism that makes transfer learning work. Let's go deeper into understanding it.
🧊 What Does "Freezing" Actually Mean?
When you freeze a layer, you set its parameters' requires_grad = False.
This tells PyTorch: "Don't compute gradients for these parameters during backpropagation."
No gradients → No updates → Weights stay exactly as they were from pretraining. The frozen knowledge is preserved perfectly.
🔬 Selective Freezing Strategies
Strategy A: Freeze All → Unfreeze Gradually
import torchvision.models as models
model = models.resnet50(weights='IMAGENET1K_V1')
def freeze_all(model):
for param in model.parameters():
param.requires_grad = False
def unfreeze_layer4_and_fc(model):
"""
Common strategy: unfreeze only the last residual block (layer4)
and the final fully connected layer.
Early layers learn general features — no need to change them.
"""
# Keep layer1, layer2, layer3 frozen
for name, param in model.named_parameters():
if 'layer4' in name or 'fc' in name:
param.requires_grad = True
print(f" UNFROZEN: {name}")
else:
param.requires_grad = False
freeze_all(model)
print("Phase 1: All frozen ❄️\n")
unfreeze_layer4_and_fc(model)
print("\nPhase 2: Last block unfrozen 🔥")
Strategy B: Progressive Unfreezing (ULMFiT Style)
Unfreeze one layer at a time, from the top downward. This is the approach pioneered by Jeremy Howard in ULMFiT and is extremely effective for NLP and vision models.
def get_resnet_layer_groups(model):
"""
Group ResNet layers from top (task-specific) to bottom (general).
We unfreeze top-to-bottom progressively.
"""
return [
model.fc, # Group 1: Classifier head (unfreeze first)
model.layer4, # Group 2: Last residual block
model.layer3, # Group 3: Third residual block
model.layer2, # Group 4: Second residual block
model.layer1, # Group 5: First residual block
model.conv1, # Group 6: Initial conv (unfreeze last)
]
def progressive_unfreeze(model, phase):
"""
Unfreeze layers progressively.
phase=1: only head
phase=2: head + layer4
phase=3: head + layer4 + layer3 ... etc.
"""
groups = get_resnet_layer_groups(model)
# First freeze everything
for param in model.parameters():
param.requires_grad = False
# Then unfreeze top `phase` groups
for group in groups[:phase]:
for param in group.parameters():
param.requires_grad = True
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Phase {phase}: {trainable:,} trainable parameters")
# Usage across training epochs:
for phase in range(1, 5):
progressive_unfreeze(model, phase)
# ... run a few training epochs here
- Unfreeze 1 new layer group every 3–5 epochs
- Halve the learning rate each time you unfreeze a new group
- Monitor validation loss — if it spikes, you unfroze too aggressively
- Early layers (edges, textures) rarely need fine-tuning for similar domains
Strategy C: Differential Learning Rates
Instead of a single learning rate for all layers, assign different learning rates to different layer groups. This is the most sophisticated and effective approach.
# Different LR for each depth of the network.
# Lower layers: much smaller LR (preserve general knowledge)
# Higher layers: larger LR (task-specific adaptation)
param_groups = [
{'params': model.layer1.parameters(), 'lr': 1e-6}, # Lowest LR
{'params': model.layer2.parameters(), 'lr': 3e-6},
{'params': model.layer3.parameters(), 'lr': 1e-5},
{'params': model.layer4.parameters(), 'lr': 3e-5},
{'params': model.fc.parameters(), 'lr': 1e-4}, # Highest LR
]
optimizer = torch.optim.AdamW(param_groups, weight_decay=1e-4)
print("Differential learning rates set:")
for i, g in enumerate(optimizer.param_groups, 1):
print(f" Group {i}: lr = {g['lr']:.1e}")
Output:
Differential learning rates set:
Group 1: lr = 1.0e-06
Group 2: lr = 3.0e-06
Group 3: lr = 1.0e-05
Group 4: lr = 3.0e-05
Group 5: lr = 1.0e-04
🗺️ Section 7: Which Strategy Should You Use? Decision Framework
Not sure whether to use Feature Extraction or Fine-Tuning? Use this decision tree every time!
↓
┌─────────────────────────────────────────────────┐
│ Q1: Is your task domain similar to ImageNet? │
│ (Natural photos, animals, objects, scenes) │
└─────────────────────────────────────────────────┘
YES ↓ NO ↓
┌───────────────────┐ ┌───────────────────────────┐
│ Q2: Data size? │ │ Different domain │
│ │ │ (Medical, Satellite, X-ray│
│ Small (<2K): │ │ Microscopy, etc.) │
│ → FEATURE EXTRACT│ │ │
│ Medium (2K–10K): │ │ Small data: Feature Extract│
│ → FINE-TUNE last │ │ Large data: FULL Fine-Tune│
│ 2–3 blocks │ └───────────────────────────┘
│ Large (>10K): │
│ → FULL FINE-TUNE │
└───────────────────┘
Similar Domain + Small Data → Feature Extraction (safest choice)
Similar Domain + Large Data → Fine-tune last few layers
Different Domain + Small Data → Feature Extraction + more augmentation
Different Domain + Large Data → Full fine-tuning (or train from scratch)
🌟 Section 8: Popular Pretrained Models
The world of pretrained models has exploded. Here's your map of the landscape.
🖼️ Vision Models
| Model | Best For | Params | Speed |
|---|---|---|---|
| ResNet50/101 | General classification, classic baseline | 25M / 44M | Fast |
| EfficientNet-B0 to B7 | Best accuracy/efficiency tradeoff | 5M – 66M | Very Fast |
| Vision Transformer (ViT) | High accuracy when data is large | 86M – 307M | Moderate |
| ConvNeXt V2 | Modern CNN, rivals ViT accuracy | 28M – 198M | Fast |
| DINOv2 (Meta) | Self-supervised, amazing features | 22M – 1.1B | Varies |
| CLIP (OpenAI) | Image + text understanding together | 63M – 428M | Moderate |
📝 NLP / Language Models
| Model | Best For | Key Strength |
|---|---|---|
| BERT / RoBERTa | Classification, NER, Q&A | Bidirectional understanding |
| GPT-2 / GPT-Neo | Text generation, completion | Autoregressive generation |
| LLaMA 3 / Mistral 7B | General LLM tasks, fine-tuning | Open source, highly fine-tunable |
| sentence-transformers | Semantic search, embeddings | Fast, great sentence vectors |
| Whisper (OpenAI) | Speech-to-text transcription | Multilingual, very accurate |
🤖 Multimodal Models
LLaVA / Qwen-VL: Vision + Language — ask questions about images
- BLIP-2: Connect any image encoder to any LLM
- Stable Diffusion (fine-tunable): Image generation with your own style
- SAM 2 (Meta): Segment anything in images AND videos
- Gemini / GPT-4V: Closed-source multimodal giants
- 🤗 Hugging Face Hub — huggingface.co/models (700,000+ models!)
- 🔦 TorchVision Models — torchvision.models (PyTorch built-in)
- 🧠 TensorFlow Hub — tfhub.dev
- 📦 timm library — 700+ vision models, easiest API
⚡ Section 9: Fine-Tuning LLMs — LoRA, QLoRA & PEFT
Fine-tuning a 7 billion parameter model sounds impossible on a regular laptop, right? Not anymore! Meet the revolution in efficient fine-tuning.
🔑 The Core Problem
Full fine-tuning of LLaMA-3 7B requires storing gradients for 7 billion parameters. That needs ~56GB of GPU memory. Most people don't have that.
Parameter-Efficient Fine-Tuning (PEFT). You only train a tiny fraction of the parameters — sometimes just 0.1% — and still get incredible results.
🛸 LoRA: Low-Rank Adaptation
LoRA is the most important fine-tuning innovation of the last few years. Here's the key idea:
Instead of updating the huge weight matrix W directly, you add two tiny matrices A and B alongside it. You only train A and B. The original W stays frozen.
LoRA adds: ΔW = A × B
A: (4096 × r) where r = rank (e.g., 8 or 16)
B: (r × 4096)
Total new params: 4096×8 + 8×4096 = 65,536 (0.4% of original!)
Final output: (W + ΔW) × input = (W + A×B) × input
# Install: pip install transformers peft bitsandbytes accelerate
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, TaskType
import torch
# --- Step 1: Load base model in 4-bit quantization (QLoRA) ---
# This reduces memory from ~28GB to ~6GB for a 7B model!
bnb_config = BitsAndBytesConfig(
load_in_4bit=True, # Quantize to 4-bit
bnb_4bit_quant_type="nf4", # NF4 quantization (best quality)
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True, # Double quantization for extra savings
)
model_name = "meta-llama/Llama-3.2-1B" # 1B for demo (use 7B for real tasks)
model = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=bnb_config,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
print(f"Base model loaded!")
# --- Step 2: Configure LoRA ---
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM, # Language modeling task
r=16, # Rank — higher = more capacity but more params
lora_alpha=32, # Scaling factor (usually 2x rank)
lora_dropout=0.05, # Dropout for regularization
target_modules=[ # Which layers to apply LoRA to
"q_proj", # Query projection (attention)
"k_proj", # Key projection
"v_proj", # Value projection
"o_proj", # Output projection
],
bias="none",
)
# --- Step 3: Wrap model with LoRA ---
peft_model = get_peft_model(model, lora_config)
# See how few params we're training!
peft_model.print_trainable_parameters()
# Output: trainable params: 3,407,872 || all params: 1,239,177,216 || trainable%: 0.2750
from transformers import TrainingArguments, Trainer, DataCollatorForLanguageModeling
from datasets import Dataset
# --- Step 4: Prepare a simple training dataset ---
# Example: fine-tune on customer support conversations
training_data = [
{"text": "User: How do I reset my password?\nAssistant: Click 'Forgot Password' on the login page and follow the email instructions."},
{"text": "User: Where is my order?\nAssistant: You can track your order using the tracking number in your confirmation email."},
# Add hundreds more examples here...
]
dataset = Dataset.from_list(training_data)
def tokenize_function(examples):
return tokenizer(
examples["text"],
truncation=True,
max_length=256,
padding="max_length"
)
tokenized_dataset = dataset.map(tokenize_function, batched=True)
# --- Step 5: Training arguments ---
training_args = TrainingArguments(
output_dir="./lora_finetuned_model",
num_train_epochs=3,
per_device_train_batch_size=4,
gradient_accumulation_steps=4, # Simulate batch_size=16
learning_rate=2e-4, # LoRA uses higher LR than full fine-tuning
fp16=True, # Mixed precision for speed
logging_steps=10,
save_steps=50,
warmup_ratio=0.05,
lr_scheduler_type="cosine",
report_to="none", # Disable wandb for simplicity
)
# Data collator handles padding automatically
data_collator = DataCollatorForLanguageModeling(tokenizer, mlm=False)
# --- Step 6: Train! ---
trainer = Trainer(
model=peft_model,
args=training_args,
train_dataset=tokenized_dataset,
data_collator=data_collator,
)
trainer.train()
print("✅ LoRA fine-tuning complete!")
# --- Step 7: Save the LoRA adapter (tiny! ~40MB vs 14GB for full model) ---
peft_model.save_pretrained("./my_lora_adapter")
print("💾 LoRA adapter saved — only ~40MB!")
Full Fine-Tuning: All params train. Needs 80GB+ GPU. Best results. Very expensive.
LoRA: Tiny adapters train. Needs ~20GB GPU. Near full quality. 10x cheaper.
QLoRA: LoRA + 4-bit quantization. Needs ~6GB GPU. Slightly lower quality. 100x cheaper.
.
🐕 Section 10: Complete End-to-End Project — Dog Breed Classifier
Let's build a real, complete project from scratch to deployment. We'll classify 10 dog breeds using EfficientNet-B2 + Fine-Tuning.
Project Structure
dog_breed_classifier/
├── data/
│ ├── train/
│ │ ├── labrador/ # 400 images
│ │ ├── poodle/ # 400 images
│ │ └── ... # 8 more breeds
│ └── val/
│ ├── labrador/ # 100 images
│ └── ...
├── train.py
├── predict.py
└── requirements.txt
Complete Training Script (train.py)
"""
Dog Breed Classifier using EfficientNet-B2 + Fine-Tuning
Complete production-ready training script.
"""
import torch
import torch.nn as nn
import torchvision.models as models
import torchvision.transforms as transforms
from torch.utils.data import DataLoader
from torchvision.datasets import ImageFolder
import torch.optim as optim
import time
import os
# ─── CONFIG ───────────────────────────────────────────────────────
CONFIG = {
'data_dir': 'data',
'num_classes': 10,
'batch_size': 32,
'phase1_epochs': 5, # Feature extraction warmup
'phase2_epochs': 20, # Fine-tuning
'phase1_lr': 1e-3,
'phase2_lr_head': 1e-4,
'phase2_lr_backbone': 1e-5,
'weight_decay': 1e-4,
'model_save': 'best_dog_classifier.pth',
'device': 'cuda' if torch.cuda.is_available() else 'cpu'
}
print(f"Using device: {CONFIG['device']}")
# ─── DATA TRANSFORMS ──────────────────────────────────────────────
MEAN = [0.485, 0.456, 0.406]
STD = [0.229, 0.224, 0.225]
train_transforms = transforms.Compose([
transforms.Resize(300),
transforms.RandomCrop(260),
transforms.RandomHorizontalFlip(p=0.5),
transforms.RandomRotation(15),
transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3),
transforms.RandomGrayscale(p=0.05),
transforms.ToTensor(),
transforms.Normalize(MEAN, STD),
])
val_transforms = transforms.Compose([
transforms.Resize((260, 260)),
transforms.ToTensor(),
transforms.Normalize(MEAN, STD),
])
# ─── DATASETS & LOADERS ───────────────────────────────────────────
train_ds = ImageFolder(os.path.join(CONFIG['data_dir'], 'train'), train_transforms)
val_ds = ImageFolder(os.path.join(CONFIG['data_dir'], 'val'), val_transforms)
train_loader = DataLoader(train_ds, batch_size=CONFIG['batch_size'],
shuffle=True, num_workers=4, pin_memory=True)
val_loader = DataLoader(val_ds, batch_size=CONFIG['batch_size'],
shuffle=False, num_workers=4, pin_memory=True)
print(f"Training samples: {len(train_ds)}")
print(f"Validation samples: {len(val_ds)}")
print(f"Classes: {train_ds.classes}")
# ─── MODEL SETUP ──────────────────────────────────────────────────
def build_model(num_classes, freeze_backbone=True):
"""Build EfficientNet-B2 with custom head"""
model = models.efficientnet_b2(weights='IMAGENET1K_V1')
if freeze_backbone:
for param in model.parameters():
param.requires_grad = False
# Replace classifier
in_features = model.classifier[1].in_features # 1408 for B2
model.classifier = nn.Sequential(
nn.Dropout(p=0.3),
nn.Linear(in_features, 512),
nn.BatchNorm1d(512),
nn.ReLU(),
nn.Dropout(p=0.2),
nn.Linear(512, num_classes)
)
return model
# ─── TRAINING UTILITIES ───────────────────────────────────────────
def run_epoch(model, loader, optimizer, criterion, device, is_train=True):
model.train() if is_train else model.eval()
total_loss, correct, total = 0.0, 0, 0
context = torch.enable_grad() if is_train else torch.no_grad()
with context:
for images, labels in loader:
images, labels = images.to(device), labels.to(device)
outputs = model(images)
loss = criterion(outputs, labels)
if is_train:
optimizer.zero_grad()
loss.backward()
optimizer.step()
total_loss += loss.item()
preds = outputs.argmax(dim=1)
correct += (preds == labels).sum().item()
total += labels.size(0)
return total_loss / len(loader), 100.0 * correct / total
# ─── PHASE 1: FEATURE EXTRACTION ──────────────────────────────────
print("\n" + "="*55)
print(" PHASE 1: Feature Extraction (Warming Up Head)")
print("="*55)
model = build_model(CONFIG['num_classes'], freeze_backbone=True)
model = model.to(CONFIG['device'])
optimizer1 = optim.Adam(
filter(lambda p: p.requires_grad, model.parameters()),
lr=CONFIG['phase1_lr']
)
criterion = nn.CrossEntropyLoss(label_smoothing=0.1)
scheduler1 = optim.lr_scheduler.OneCycleLR(
optimizer1, max_lr=CONFIG['phase1_lr'],
steps_per_epoch=len(train_loader),
epochs=CONFIG['phase1_epochs']
)
best_val_acc = 0.0
for epoch in range(1, CONFIG['phase1_epochs'] + 1):
t0 = time.time()
tr_loss, tr_acc = run_epoch(model, train_loader, optimizer1, criterion, CONFIG['device'])
vl_loss, vl_acc = run_epoch(model, val_loader, None, criterion, CONFIG['device'], False)
scheduler1.step()
elapsed = time.time() - t0
print(f"Ep {epoch:02d}/{CONFIG['phase1_epochs']} | "
f"Train {tr_acc:.1f}% | Val {vl_acc:.1f}% | {elapsed:.1f}s")
if vl_acc > best_val_acc:
best_val_acc = vl_acc
torch.save(model.state_dict(), CONFIG['model_save'])
print(f" ✅ Saved! Best val: {best_val_acc:.1f}%")
# ─── PHASE 2: FINE-TUNING ─────────────────────────────────────────
print("\n" + "="*55)
print(" PHASE 2: Fine-Tuning (Adapting Backbone)")
print("="*55)
# Load best weights from phase 1
model.load_state_dict(torch.load(CONFIG['model_save']))
# Unfreeze last 2 feature blocks + classifier
for param in model.parameters():
param.requires_grad = False # Re-freeze all
# Selectively unfreeze deep blocks
blocks = list(model.features.children())
for block in blocks[6:]: # Unfreeze last 2 blocks (6, 7)
for param in block.parameters():
param.requires_grad = True
for param in model.classifier.parameters():
param.requires_grad = True
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Fine-tuning {trainable:,} parameters")
optimizer2 = optim.AdamW([
{'params': model.features[6:].parameters(),
'lr': CONFIG['phase2_lr_backbone']},
{'params': model.classifier.parameters(),
'lr': CONFIG['phase2_lr_head']},
], weight_decay=CONFIG['weight_decay'])
scheduler2 = optim.lr_scheduler.CosineAnnealingLR(
optimizer2, T_max=CONFIG['phase2_epochs'], eta_min=1e-7
)
for epoch in range(1, CONFIG['phase2_epochs'] + 1):
t0 = time.time()
tr_loss, tr_acc = run_epoch(model, train_loader, optimizer2, criterion, CONFIG['device'])
vl_loss, vl_acc = run_epoch(model, val_loader, None, criterion, CONFIG['device'], False)
scheduler2.step()
elapsed = time.time() - t0
print(f"Ep {epoch:02d}/{CONFIG['phase2_epochs']} | "
f"Train {tr_acc:.1f}% | Val {vl_acc:.1f}% | {elapsed:.1f}s")
if vl_acc > best_val_acc:
best_val_acc = vl_acc
torch.save(model.state_dict(), CONFIG['model_save'])
print(f" 🏆 New Best! Val: {best_val_acc:.1f}%")
print(f"\n✅ Training complete! Best Val Accuracy: {best_val_acc:.1f}%")
Prediction Script (predict.py)
"""
predict.py — Load fine-tuned model and classify a dog image.
"""
import torch
import torchvision.models as models
import torchvision.transforms as transforms
from PIL import Image
import torch.nn as nn
CLASS_NAMES = [
'beagle', 'bulldog', 'chihuahua', 'doberman', 'golden_retriever',
'husky', 'labrador', 'poodle', 'rottweiler', 'yorkie'
]
MEAN = [0.485, 0.456, 0.406]
STD = [0.229, 0.224, 0.225]
def load_model(weights_path, num_classes=10):
model = models.efficientnet_b2(weights=None) # No pretrained this time
in_features = model.classifier[1].in_features
model.classifier = nn.Sequential(
nn.Dropout(p=0.3),
nn.Linear(in_features, 512),
nn.BatchNorm1d(512),
nn.ReLU(),
nn.Dropout(p=0.2),
nn.Linear(512, num_classes)
)
model.load_state_dict(torch.load(weights_path, map_location='cpu'))
model.eval()
return model
def predict(image_path, model):
transform = transforms.Compose([
transforms.Resize((260, 260)),
transforms.ToTensor(),
transforms.Normalize(MEAN, STD),
])
image = Image.open(image_path).convert('RGB')
tensor = transform(image).unsqueeze(0) # Add batch dimension
with torch.no_grad():
outputs = model(tensor)
probs = torch.softmax(outputs, dim=1)
top5 = torch.topk(probs, 5)
print(f"\n🐕 Predictions for: {image_path}")
print("-" * 35)
for prob, idx in zip(top5.values[0], top5.indices[0]):
bar = "█" * int(prob.item() * 30)
print(f" {CLASS_NAMES[idx]:20s} {bar} {prob.item()*100:.1f}%")
predicted_class = CLASS_NAMES[top5.indices[0][0].item()]
confidence = top5.values[0][0].item() * 100
print(f"\n✅ Result: {predicted_class} ({confidence:.1f}% confidence)")
return predicted_class
# Run prediction
model = load_model('best_dog_classifier.pth')
result = predict('test_dog.jpg', model)
Sample Output:
🐕 Predictions for: test_dog.jpg
-----------------------------------
labrador ██████████████████████ 76.3%
golden_retriever █████ 15.2%
beagle ██ 5.1%
bulldog █ 2.1%
husky 1.3%
✅ Result: labrador (76.3% confidence)
🚫 Section 11: Common Mistakes & How to Avoid Them
Every pretrained model expects input normalized with the same mean/std used during pretraining. For ImageNet models, always use:
mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]Using random normalization values will give terrible results — the model "sees" completely wrong colors and scales.
Never use a learning rate of 1e-3 or higher when fine-tuning pretrained layers. This destroys the carefully learned weights in just a few batches. Rule of thumb: Phase 2 LR should be 10x–100x smaller than Phase 1 LR.
Jumping straight to fine-tuning without first warming up the new head causes large random gradients to backpropagate through the pretrained layers, instantly wrecking them. Always warm up the head first (Phase 1)!
BatchNorm and Dropout behave differently in training vs evaluation mode. Always call
model.eval() before inference and model.train()
before training. Missing this causes inconsistent and wrong predictions.
Fine-tuning ViT-Large (307M params) on 500 images = massive overfitting. Match your model size to your data size. For <2000 samples: use EfficientNet-B0 or ResNet18. Larger models need more data to benefit from fine-tuning.
- ✅ Always use the correct preprocessing (normalization) for your pretrained model
- ✅ Warm up the head first (Phase 1) before fine-tuning the backbone (Phase 2)
- ✅ Use differential learning rates — lower for earlier layers
- ✅ Use label smoothing (0.05–0.1) to prevent overconfident predictions
- ✅ Use data augmentation generously — especially when data is small
- ✅ Monitor both training and validation loss — catch overfitting early
- ✅ Save the best checkpoint (val acc), not the last one
- ✅ Use cosine LR scheduling for fine-tuning
- ✅ Try
timmlibrary for the easiest pretrained model API
🏆 Hero-Level Summary: Everything You've Mastered!
- 🔹 Transfer Learning = Borrow knowledge from massive pretrained models. Stand on giants' shoulders.
- 🔹 Feature Extraction = Freeze all pretrained weights. Train only a small new head. Fast, works with tiny data.
- 🔹 Fine-Tuning = Unfreeze some/all layers. Let the model adapt. Needs more data. Gets better accuracy.
- 🔹 Two-Phase Strategy = Warm up head first → then fine-tune backbone. Industry gold standard.
- 🔹 Differential LR = Different learning rates for different layer depths. Lower LR for early layers.
- 🔹 Progressive Unfreezing = Unfreeze layer by layer, top to bottom. Gentlest and most effective.
- 🔹 Top Vision Models = EfficientNet, ConvNeXt V2, ViT, DINOv2, CLIP.
- 🔹 LoRA / QLoRA = Fine-tune 7B LLMs on a single GPU by training tiny adapter matrices.
- 🔹 PEFT = Family of efficient fine-tuning methods. LoRA is the most popular member.
- 🔹 Always normalize correctly! Wrong normalization ruins everything.
Keep experimenting, keep building, and keep learning! 🐼✨
Comments
Post a Comment