Skip to main content

Pretrained Models: Feature Extraction & Fine-Tuning

Calculating read time…

Imagine you want to become a chef. You have two choices:

Option A: Start from scratch. Learn what fire is. Learn how to hold a knife. Learn every spice from zero. Takes years!

Option B: Find a world-class chef who already knows everything. Ask them to teach you just the last part — your specific cuisine. Done in weeks!




💡 That's exactly what Pretrained Models do for your AI!

Instead of training a neural network from zero (which takes weeks and massive compute), you borrow a model that was already trained by Google, Meta, or OpenAI on millions of images or texts. Then you either use it directly or teach it your specific task. This is called Transfer Learning — and it's the #1 superpower of modern Deep Learning. ⚡

🧠 Section 1: What is Transfer Learning? (The Foundation)

Every time a neural network trains on data, it doesn't just memorize answers. It builds internal knowledge — patterns, shapes, textures, relationships — stored inside its millions of weight parameters.

Transfer Learning is the idea of taking that stored knowledge and moving it to a new problem. You're not starting from zero. You're standing on the shoulders of giants. 🏔️

🧩 A Simple Analogy: The Expert Photographer

Say you're a professional wildlife photographer for 10 years. You've mastered lighting, composition, timing, and camera settings.

Now your friend asks you to photograph their wedding. You don't need to re-learn what a camera is! You transfer all your existing skills and just learn the specific style of wedding photography. Done fast. Done well.

Neural networks do the same thing. A model trained on 14 million images has learned to recognize edges, colors, textures, faces, objects. These are universal visual skills. You can borrow them for your specific task — whether it's detecting skin diseases or counting cars.


🏗️ Section 2: What Are Pretrained Models? How Were They Built?

A pretrained model is a neural network that has already been trained on a massive dataset — usually by a research lab or big tech company — and whose learned weights are made available to the public.

📦 What's Inside a Pretrained Model?

Think of a deep neural network like a multi-story building. Each floor (layer) is responsible for learning something different:

INPUT IMAGE → [Floor 1: Edges & Lines] → [Floor 2: Corners & Curves]
→ [Floor 3: Textures & Patterns] → [Floor 4: Parts (eyes, wheels, leaves)]
→ [Floor 5: Objects (face, car, tree)] → [Roof: Final Classification]

The lower floors learn general, universal features that are useful for any image task. The higher floors learn task-specific features that only matter for the original training task (e.g., identifying ImageNet classes).

When you borrow a pretrained model, you get ALL these floors for free! You only need to decide what to do with the roof.

🌍 Where Do These Models Come From?

  • ImageNet Challenge (ILSVRC): Models trained on 1.2 million images across 1,000 categories. Gave us ResNet, VGG, Inception, EfficientNet.
  • Google's JFT-300M / JFT-3B: 3 billion images! Used to train Vision Transformers (ViT) — the current state of the art.
  • Common Crawl (for NLP): Petabytes of web text used to train BERT, GPT, LLaMA, Mistral and other language models.
  • LAION-5B (for Multimodal): 5 billion image-text pairs used to train CLIP, DALL-E, Stable Diffusion.
💡 Fun Fact:
Training GPT-4 from scratch is estimated to cost over $100 million in compute. Fine-tuning it for your task might cost $50–$500 on cloud GPUs. Transfer learning is not just convenient — it's economically transformative! 💰

⚖️ Section 3: The Two Big Strategies

Once you have a pretrained model, you have two main ways to use it. Understanding the difference between them is the most important thing in this guide.

Strategy 1: Feature Extraction 🧲

You use the pretrained model as a fixed feature extractor. You freeze ALL its weights — nothing changes inside the pretrained model. You only add and train a small new "head" (classifier) on top.

┌─────────────────────────────────────────┐
│ PRETRAINED MODEL (ALL FROZEN ❄️) │
│ Conv1 → Conv2 → Conv3 → ... → Pool │
│ (weights do NOT change at all) │
└─────────────────────────────────────────┘
                  ↓
      Feature Vector (e.g., 2048 numbers)
                  ↓
┌─────────────────────────────────────────┐
│ YOUR NEW HEAD (Trainable 🔥) │
│ Dense(256) → Dropout → Dense(N classes)│
└─────────────────────────────────────────┘

Think of it like: You hire a brilliant analyst (the pretrained model) who generates a detailed report (feature vector) on every piece of data. You then train a simple decision-maker on those reports. The analyst never changes. You only train the decision-maker.

Strategy 2: Fine-Tuning 🎛️

You unfreeze some or all layers of the pretrained model and let them update their weights on your new data — but with a very small learning rate so you don't destroy the existing knowledge.

┌─────────────────────────────────────────┐
│ PRETRAINED MODEL │
│ Conv1 → Conv2 (FROZEN ❄️) │
│ Conv3 → Conv4 → Conv5 (TRAINABLE 🔥) │
└─────────────────────────────────────────┘
                  ↓
┌─────────────────────────────────────────┐
│ YOUR NEW HEAD (Trainable 🔥) │
│ Dense(256) → Dropout → Dense(N classes)│
└─────────────────────────────────────────┘

Think of it like: Now you let the expert analyst actually adjust their analysis style to better fit your specific domain. They don't start from scratch — they gently adapt their deep expertise to your new context.

📊 Feature Extraction vs Fine-Tuning: Side-by-Side

Aspect Feature Extraction Fine-Tuning
Pretrained weights Completely frozen ❄️ Partially or fully unfrozen 🔥
Training speed Very fast ⚡ Slower (more params)
Data needed Works with very little (100–1000 samples) Needs more (1000–10000+ samples)
Risk of overfitting Lower Higher (if data is small)
Final accuracy Good Usually better 🏆
GPU memory Lower Higher
Best when... Small dataset, similar domain Larger dataset, different domain

🧲 Section 4: Feature Extraction — Deep Dive with Code

Let's build a real image classifier using Feature Extraction with ResNet50 (a classic, battle-tested pretrained model).

Step 1: Load the Pretrained Model

import torch
import torch.nn as nn
import torchvision.models as models
import torchvision.transforms as transforms
from torch.utils.data import DataLoader
from torchvision.datasets import ImageFolder

# Load ResNet50 pretrained on ImageNet
# weights='IMAGENET1K_V1' loads the classic pretrained weights
model = models.resnet50(weights='IMAGENET1K_V1')

print("ResNet50 loaded!")
print(f"Total parameters: {sum(p.numel() for p in model.parameters()):,}")

Output:

ResNet50 loaded!
Total parameters: 25,557,032

Step 2: Freeze ALL Pretrained Layers

# Freeze every single layer in the pretrained model.
# requires_grad=False means: "Don't update these weights during training"
for param in model.parameters():
    param.requires_grad = False

# Verify: count trainable parameters (should be 0 now)
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Trainable parameters after freezing: {trainable:,}")
# Output: Trainable parameters after freezing: 0

Step 3: Replace the Final Classification Head

# ResNet50's final layer is model.fc
# By default it outputs 1000 classes (ImageNet).
# We replace it with our own classifier for OUR number of classes.

NUM_CLASSES = 10  # Example: 10 dog breeds

# The original model.fc takes 2048 features as input
in_features = model.fc.in_features  # = 2048

# Build a custom classification head
# This head IS trainable — only these params will update!
model.fc = nn.Sequential(
    nn.Linear(in_features, 256),
    nn.ReLU(),
    nn.Dropout(p=0.4),
    nn.Linear(256, NUM_CLASSES)
)

# Verify: now only the new head is trainable
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total     = sum(p.numel() for p in model.parameters())

print(f"Trainable params: {trainable:,}")
print(f"Total params:     {total:,}")
print(f"Training only {trainable/total*100:.2f}% of the model!")

Output:

Trainable params: 524,810
Total parameters: 26,081,842
Training only 2.01% of the model!
✅ This is the beauty of Feature Extraction!
You're training just 2% of the total parameters. Training is lightning fast and almost impossible to overfit, even on very small datasets. 🏎️

Step 4: Prepare Data with Correct Transforms

# CRITICAL: ImageNet pretrained models expect specific input preprocessing.
# Always normalize with ImageNet mean and std — the model was trained with these!

IMAGENET_MEAN = [0.485, 0.456, 0.406]
IMAGENET_STD  = [0.229, 0.224, 0.225]

train_transforms = transforms.Compose([
    transforms.Resize((256, 256)),
    transforms.RandomCrop(224),           # ResNet expects 224x224
    transforms.RandomHorizontalFlip(),    # Basic augmentation
    transforms.ColorJitter(brightness=0.2, contrast=0.2),
    transforms.ToTensor(),
    transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD)
])

val_transforms = transforms.Compose([
    transforms.Resize((224, 224)),        # No augmentation for validation
    transforms.ToTensor(),
    transforms.Normalize(mean=IMAGENET_MEAN, std=IMAGENET_STD)
])

# Load your dataset (folder structure: data/train/class1, data/train/class2, ...)
train_dataset = ImageFolder('data/train', transform=train_transforms)
val_dataset   = ImageFolder('data/val',   transform=val_transforms)

train_loader  = DataLoader(train_dataset, batch_size=32, shuffle=True,  num_workers=4)
val_loader    = DataLoader(val_dataset,   batch_size=32, shuffle=False, num_workers=4)

print(f"Training samples:   {len(train_dataset)}")
print(f"Validation samples: {len(val_dataset)}")
print(f"Classes:            {train_dataset.classes}")

Step 5: Train ONLY the New Head

import torch.optim as optim

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model  = model.to(device)

# Only pass trainable parameters to the optimizer
# This is important! Only model.fc params are trainable.
optimizer = optim.Adam(
    filter(lambda p: p.requires_grad, model.parameters()),
    lr=1e-3
)

criterion  = nn.CrossEntropyLoss()
scheduler  = optim.lr_scheduler.StepLR(optimizer, step_size=5, gamma=0.5)


def train_one_epoch(model, loader, optimizer, criterion, device):
    model.train()
    total_loss, correct, total = 0, 0, 0

    for images, labels in loader:
        images, labels = images.to(device), labels.to(device)

        optimizer.zero_grad()
        outputs = model(images)
        loss    = criterion(outputs, labels)
        loss.backward()
        optimizer.step()

        total_loss += loss.item()
        preds       = outputs.argmax(dim=1)
        correct    += (preds == labels).sum().item()
        total      += labels.size(0)

    return total_loss / len(loader), correct / total * 100


def evaluate(model, loader, criterion, device):
    model.eval()
    total_loss, correct, total = 0, 0, 0

    with torch.no_grad():
        for images, labels in loader:
            images, labels = images.to(device), labels.to(device)
            outputs = model(images)
            loss    = criterion(outputs, labels)

            total_loss += loss.item()
            preds       = outputs.argmax(dim=1)
            correct    += (preds == labels).sum().item()
            total      += labels.size(0)

    return total_loss / len(loader), correct / total * 100


# Training loop
EPOCHS = 10
best_val_acc = 0.0

print(f"Training on: {device}")
print("="*55)

for epoch in range(1, EPOCHS + 1):
    train_loss, train_acc = train_one_epoch(
        model, train_loader, optimizer, criterion, device
    )
    val_loss, val_acc = evaluate(model, val_loader, criterion, device)
    scheduler.step()

    print(f"Epoch {epoch:02d}/{EPOCHS} | "
          f"Train Loss: {train_loss:.4f} | Train Acc: {train_acc:.2f}% | "
          f"Val Loss: {val_loss:.4f} | Val Acc: {val_acc:.2f}%")

    # Save the best model
    if val_acc > best_val_acc:
        best_val_acc = val_acc
        torch.save(model.state_dict(), 'best_model_feature_extraction.pth')
        print(f"   ✅ New best model saved! (Val Acc: {val_acc:.2f}%)")

print(f"\n🏆 Best Validation Accuracy: {best_val_acc:.2f}%")

🎛️ Section 5: Fine-Tuning — Deep Dive with Code

Fine-tuning takes feature extraction one step further. After the new head has warmed up, you gently unfreeze deeper layers and let the whole model adapt to your data.

💡 The Two-Phase Strategy (Industry Best Practice):

Phase 1 (Feature Extraction): Train ONLY the new head for a few epochs. This "warms up" the head so it produces reasonable outputs.

Phase 2 (Fine-Tuning): Unfreeze some or all of the pretrained layers. Continue training with a MUCH smaller learning rate. Now the whole network adapts together.

Why two phases? If you fine-tune from the very start, the large gradients from the random new head can destroy the pretrained weights. Warm up first — then unfreeze!

Phase 1: Warm Up the Head (Same as Feature Extraction)

import torch
import torch.nn as nn
import torchvision.models as models
import torch.optim as optim

# Load pretrained EfficientNet-B0 (excellent for efficiency)
model = models.efficientnet_b0(weights='IMAGENET1K_V1')

# --- PHASE 1: Freeze everything ---
for param in model.parameters():
    param.requires_grad = False

# Replace classifier head
# EfficientNet's classifier is model.classifier[1]
in_features = model.classifier[1].in_features  # = 1280

NUM_CLASSES = 5

model.classifier = nn.Sequential(
    nn.Dropout(p=0.3, inplace=True),
    nn.Linear(in_features, 256),
    nn.ReLU(),
    nn.Dropout(p=0.3),
    nn.Linear(256, NUM_CLASSES)
)

device    = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model     = model.to(device)

# Phase 1 optimizer — high learning rate is OK since only head trains
optimizer_phase1 = optim.Adam(
    filter(lambda p: p.requires_grad, model.parameters()),
    lr=1e-3
)
criterion = nn.CrossEntropyLoss()

print("Phase 1: Training new head only...")
# (Run train_one_epoch for 5 epochs here — same function as above)
# ... (training loop same as feature extraction section)
print("Phase 1 complete! Head is warmed up ✅")

Phase 2: Unfreeze & Fine-Tune

# --- PHASE 2: Selectively unfreeze layers ---

# Strategy: unfreeze ONLY the last few blocks of EfficientNet
# This gives us the benefits of fine-tuning without catastrophic forgetting

# First unfreeze everything
for param in model.parameters():
    param.requires_grad = True

# Then re-freeze the early layers (feature 0 through 5 = early blocks)
# We keep features 6, 7, 8 (deeper blocks) + classifier trainable
layers_to_freeze = list(model.features.children())[:6]  # First 6 blocks

for layer in layers_to_freeze:
    for param in layer.parameters():
        param.requires_grad = False

# Count trainable parameters now
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total     = sum(p.numel() for p in model.parameters())
print(f"Phase 2: Training {trainable:,} / {total:,} parameters")
print(f"         ({trainable/total*100:.1f}% of model)")

# CRITICAL: Use a much smaller learning rate for fine-tuning!
# Large LR will destroy pretrained knowledge. Rule of thumb: 10x–100x smaller.
optimizer_phase2 = optim.Adam([
    # Deeper pretrained layers — very small LR
    {'params': model.features[6:].parameters(), 'lr': 1e-5},
    # New classifier head — slightly larger LR
    {'params': model.classifier.parameters(),   'lr': 1e-4},
])

# Cosine annealing LR scheduler — standard in fine-tuning
scheduler_phase2 = optim.lr_scheduler.CosineAnnealingLR(
    optimizer_phase2, T_max=15, eta_min=1e-7
)

print("\nPhase 2: Fine-tuning with differential learning rates...")

EPOCHS_PHASE2 = 15
best_val_acc   = 0.0

for epoch in range(1, EPOCHS_PHASE2 + 1):
    train_loss, train_acc = train_one_epoch(
        model, train_loader, optimizer_phase2, criterion, device
    )
    val_loss, val_acc = evaluate(model, val_loader, criterion, device)
    scheduler_phase2.step()

    print(f"Epoch {epoch:02d}/{EPOCHS_PHASE2} | "
          f"Train: {train_acc:.2f}% | Val: {val_acc:.2f}%")

    if val_acc > best_val_acc:
        best_val_acc = val_acc
        torch.save(model.state_dict(), 'best_model_finetuned.pth')
        print(f"   ✅ New best: {val_acc:.2f}%")

print(f"\n🏆 Fine-tuned Best Val Accuracy: {best_val_acc:.2f}%")

❄️ Section 6: Layer Freezing — What, Why, and How

Layer freezing is the core mechanism that makes transfer learning work. Let's go deeper into understanding it.

🧊 What Does "Freezing" Actually Mean?

When you freeze a layer, you set its parameters' requires_grad = False. This tells PyTorch: "Don't compute gradients for these parameters during backpropagation."

No gradients → No updates → Weights stay exactly as they were from pretraining. The frozen knowledge is preserved perfectly.

🔬 Selective Freezing Strategies

Strategy A: Freeze All → Unfreeze Gradually

import torchvision.models as models

model = models.resnet50(weights='IMAGENET1K_V1')

def freeze_all(model):
    for param in model.parameters():
        param.requires_grad = False

def unfreeze_layer4_and_fc(model):
    """
    Common strategy: unfreeze only the last residual block (layer4)
    and the final fully connected layer.
    Early layers learn general features — no need to change them.
    """
    # Keep layer1, layer2, layer3 frozen
    for name, param in model.named_parameters():
        if 'layer4' in name or 'fc' in name:
            param.requires_grad = True
            print(f"  UNFROZEN: {name}")
        else:
            param.requires_grad = False

freeze_all(model)
print("Phase 1: All frozen ❄️\n")
unfreeze_layer4_and_fc(model)
print("\nPhase 2: Last block unfrozen 🔥")

Strategy B: Progressive Unfreezing (ULMFiT Style)

Unfreeze one layer at a time, from the top downward. This is the approach pioneered by Jeremy Howard in ULMFiT and is extremely effective for NLP and vision models.

def get_resnet_layer_groups(model):
    """
    Group ResNet layers from top (task-specific) to bottom (general).
    We unfreeze top-to-bottom progressively.
    """
    return [
        model.fc,          # Group 1: Classifier head (unfreeze first)
        model.layer4,      # Group 2: Last residual block
        model.layer3,      # Group 3: Third residual block
        model.layer2,      # Group 4: Second residual block
        model.layer1,      # Group 5: First residual block
        model.conv1,       # Group 6: Initial conv (unfreeze last)
    ]


def progressive_unfreeze(model, phase):
    """
    Unfreeze layers progressively.
    phase=1: only head
    phase=2: head + layer4
    phase=3: head + layer4 + layer3 ... etc.
    """
    groups = get_resnet_layer_groups(model)

    # First freeze everything
    for param in model.parameters():
        param.requires_grad = False

    # Then unfreeze top `phase` groups
    for group in groups[:phase]:
        for param in group.parameters():
            param.requires_grad = True

    trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
    print(f"Phase {phase}: {trainable:,} trainable parameters")


# Usage across training epochs:
for phase in range(1, 5):
    progressive_unfreeze(model, phase)
    # ... run a few training epochs here
✅ Progressive Unfreezing Rule of Thumb:
  • Unfreeze 1 new layer group every 3–5 epochs
  • Halve the learning rate each time you unfreeze a new group
  • Monitor validation loss — if it spikes, you unfroze too aggressively
  • Early layers (edges, textures) rarely need fine-tuning for similar domains

Strategy C: Differential Learning Rates

Instead of a single learning rate for all layers, assign different learning rates to different layer groups. This is the most sophisticated and effective approach.

# Different LR for each depth of the network.
# Lower layers: much smaller LR (preserve general knowledge)
# Higher layers: larger LR (task-specific adaptation)

param_groups = [
    {'params': model.layer1.parameters(), 'lr': 1e-6},  # Lowest LR
    {'params': model.layer2.parameters(), 'lr': 3e-6},
    {'params': model.layer3.parameters(), 'lr': 1e-5},
    {'params': model.layer4.parameters(), 'lr': 3e-5},
    {'params': model.fc.parameters(),     'lr': 1e-4},  # Highest LR
]

optimizer = torch.optim.AdamW(param_groups, weight_decay=1e-4)

print("Differential learning rates set:")
for i, g in enumerate(optimizer.param_groups, 1):
    print(f"  Group {i}: lr = {g['lr']:.1e}")

Output:

Differential learning rates set:
  Group 1: lr = 1.0e-06
  Group 2: lr = 3.0e-06
  Group 3: lr = 1.0e-05
  Group 4: lr = 3.0e-05
  Group 5: lr = 1.0e-04

🗺️ Section 7: Which Strategy Should You Use? Decision Framework

Not sure whether to use Feature Extraction or Fine-Tuning? Use this decision tree every time!

START: What's your situation?
            ↓
┌─────────────────────────────────────────────────┐
│ Q1: Is your task domain similar to ImageNet? │
│ (Natural photos, animals, objects, scenes) │
└─────────────────────────────────────────────────┘
    YES ↓                           NO ↓
┌───────────────────┐       ┌───────────────────────────┐
│ Q2: Data size? │       │ Different domain │
│ │       │ (Medical, Satellite, X-ray│
│ Small (<2K): │       │ Microscopy, etc.) │
│ → FEATURE EXTRACT│       │ │
│ Medium (2K–10K): │       │ Small data: Feature Extract│
│ → FINE-TUNE last │       │ Large data: FULL Fine-Tune│
│ 2–3 blocks │       └───────────────────────────┘
│ Large (>10K): │
│ → FULL FINE-TUNE │
└───────────────────┘
💡 The 2x2 Matrix (Quick Reference):

Similar Domain + Small Data → Feature Extraction (safest choice)
Similar Domain + Large Data → Fine-tune last few layers
Different Domain + Small Data → Feature Extraction + more augmentation
Different Domain + Large Data → Full fine-tuning (or train from scratch)

🌟 Section 8: Popular Pretrained Models

The world of pretrained models has exploded. Here's your map of the landscape.

🖼️ Vision Models

Model Best For Params Speed
ResNet50/101 General classification, classic baseline 25M / 44M Fast
EfficientNet-B0 to B7 Best accuracy/efficiency tradeoff 5M – 66M Very Fast
Vision Transformer (ViT) High accuracy when data is large 86M – 307M Moderate
ConvNeXt V2 Modern CNN, rivals ViT accuracy 28M – 198M Fast
DINOv2 (Meta) Self-supervised, amazing features 22M – 1.1B Varies
CLIP (OpenAI) Image + text understanding together 63M – 428M Moderate

📝 NLP / Language Models

Model Best For Key Strength
BERT / RoBERTa Classification, NER, Q&A Bidirectional understanding
GPT-2 / GPT-Neo Text generation, completion Autoregressive generation
LLaMA 3 / Mistral 7B General LLM tasks, fine-tuning Open source, highly fine-tunable
sentence-transformers Semantic search, embeddings Fast, great sentence vectors
Whisper (OpenAI) Speech-to-text transcription Multilingual, very accurate

🤖 Multimodal Models 

LLaVA / Qwen-VL: Vision + Language — ask questions about images

  • BLIP-2: Connect any image encoder to any LLM
  • Stable Diffusion (fine-tunable): Image generation with your own style
  • SAM 2 (Meta): Segment anything in images AND videos
  • Gemini / GPT-4V: Closed-source multimodal giants
✅ Where to Find All Pretrained Models:
  • 🤗 Hugging Face Hub — huggingface.co/models (700,000+ models!)
  • 🔦 TorchVision Models — torchvision.models (PyTorch built-in)
  • 🧠 TensorFlow Hub — tfhub.dev
  • 📦 timm library — 700+ vision models, easiest API

⚡ Section 9: Fine-Tuning LLMs — LoRA, QLoRA & PEFT

Fine-tuning a 7 billion parameter model sounds impossible on a regular laptop, right? Not anymore! Meet the revolution in efficient fine-tuning.

🔑 The Core Problem

Full fine-tuning of LLaMA-3 7B requires storing gradients for 7 billion parameters. That needs ~56GB of GPU memory. Most people don't have that.

Parameter-Efficient Fine-Tuning (PEFT). You only train a tiny fraction of the parameters — sometimes just 0.1% — and still get incredible results.

🛸 LoRA: Low-Rank Adaptation

LoRA is the most important fine-tuning innovation of the last few years. Here's the key idea:

Instead of updating the huge weight matrix W directly, you add two tiny matrices A and B alongside it. You only train A and B. The original W stays frozen.

Original: W (e.g., 4096 × 4096 = 16.7M params) → FROZEN ❄️

LoRA adds: ΔW = A × B
  A: (4096 × r) where r = rank (e.g., 8 or 16)
  B: (r × 4096)
  Total new params: 4096×8 + 8×4096 = 65,536 (0.4% of original!)

Final output: (W + ΔW) × input = (W + A×B) × input
# Install: pip install transformers peft bitsandbytes accelerate

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, TaskType
import torch

# --- Step 1: Load base model in 4-bit quantization (QLoRA) ---
# This reduces memory from ~28GB to ~6GB for a 7B model!

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,               # Quantize to 4-bit
    bnb_4bit_quant_type="nf4",       # NF4 quantization (best quality)
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,  # Double quantization for extra savings
)

model_name = "meta-llama/Llama-3.2-1B"  # 1B for demo (use 7B for real tasks)

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    device_map="auto"
)

tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token

print(f"Base model loaded!")

# --- Step 2: Configure LoRA ---
lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,  # Language modeling task
    r=16,                           # Rank — higher = more capacity but more params
    lora_alpha=32,                  # Scaling factor (usually 2x rank)
    lora_dropout=0.05,              # Dropout for regularization
    target_modules=[                # Which layers to apply LoRA to
        "q_proj",                   # Query projection (attention)
        "k_proj",                   # Key projection
        "v_proj",                   # Value projection
        "o_proj",                   # Output projection
    ],
    bias="none",
)

# --- Step 3: Wrap model with LoRA ---
peft_model = get_peft_model(model, lora_config)

# See how few params we're training!
peft_model.print_trainable_parameters()
# Output: trainable params: 3,407,872 || all params: 1,239,177,216 || trainable%: 0.2750
from transformers import TrainingArguments, Trainer, DataCollatorForLanguageModeling
from datasets import Dataset

# --- Step 4: Prepare a simple training dataset ---
# Example: fine-tune on customer support conversations

training_data = [
    {"text": "User: How do I reset my password?\nAssistant: Click 'Forgot Password' on the login page and follow the email instructions."},
    {"text": "User: Where is my order?\nAssistant: You can track your order using the tracking number in your confirmation email."},
    # Add hundreds more examples here...
]

dataset = Dataset.from_list(training_data)

def tokenize_function(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        max_length=256,
        padding="max_length"
    )

tokenized_dataset = dataset.map(tokenize_function, batched=True)

# --- Step 5: Training arguments ---
training_args = TrainingArguments(
    output_dir="./lora_finetuned_model",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,   # Simulate batch_size=16
    learning_rate=2e-4,              # LoRA uses higher LR than full fine-tuning
    fp16=True,                       # Mixed precision for speed
    logging_steps=10,
    save_steps=50,
    warmup_ratio=0.05,
    lr_scheduler_type="cosine",
    report_to="none",                # Disable wandb for simplicity
)

# Data collator handles padding automatically
data_collator = DataCollatorForLanguageModeling(tokenizer, mlm=False)

# --- Step 6: Train! ---
trainer = Trainer(
    model=peft_model,
    args=training_args,
    train_dataset=tokenized_dataset,
    data_collator=data_collator,
)

trainer.train()
print("✅ LoRA fine-tuning complete!")

# --- Step 7: Save the LoRA adapter (tiny! ~40MB vs 14GB for full model) ---
peft_model.save_pretrained("./my_lora_adapter")
print("💾 LoRA adapter saved — only ~40MB!")
💡 LoRA vs QLoRA vs Full Fine-Tuning — Quick Comparison:

Full Fine-Tuning: All params train. Needs 80GB+ GPU. Best results. Very expensive.
LoRA: Tiny adapters train. Needs ~20GB GPU. Near full quality. 10x cheaper.
QLoRA: LoRA + 4-bit quantization. Needs ~6GB GPU. Slightly lower quality. 100x cheaper.

.

🐕 Section 10: Complete End-to-End Project — Dog Breed Classifier

Let's build a real, complete project from scratch to deployment. We'll classify 10 dog breeds using EfficientNet-B2 + Fine-Tuning.

Project Structure

dog_breed_classifier/
├── data/
│   ├── train/
│   │   ├── labrador/       # 400 images
│   │   ├── poodle/         # 400 images
│   │   └── ...             # 8 more breeds
│   └── val/
│       ├── labrador/       # 100 images
│       └── ...
├── train.py
├── predict.py
└── requirements.txt

Complete Training Script (train.py)

"""
Dog Breed Classifier using EfficientNet-B2 + Fine-Tuning
Complete production-ready training script.
"""

import torch
import torch.nn as nn
import torchvision.models as models
import torchvision.transforms as transforms
from torch.utils.data import DataLoader
from torchvision.datasets import ImageFolder
import torch.optim as optim
import time
import os

# ─── CONFIG ───────────────────────────────────────────────────────
CONFIG = {
    'data_dir':     'data',
    'num_classes':  10,
    'batch_size':   32,
    'phase1_epochs': 5,      # Feature extraction warmup
    'phase2_epochs': 20,     # Fine-tuning
    'phase1_lr':    1e-3,
    'phase2_lr_head': 1e-4,
    'phase2_lr_backbone': 1e-5,
    'weight_decay': 1e-4,
    'model_save':   'best_dog_classifier.pth',
    'device':       'cuda' if torch.cuda.is_available() else 'cpu'
}

print(f"Using device: {CONFIG['device']}")

# ─── DATA TRANSFORMS ──────────────────────────────────────────────
MEAN = [0.485, 0.456, 0.406]
STD  = [0.229, 0.224, 0.225]

train_transforms = transforms.Compose([
    transforms.Resize(300),
    transforms.RandomCrop(260),
    transforms.RandomHorizontalFlip(p=0.5),
    transforms.RandomRotation(15),
    transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3),
    transforms.RandomGrayscale(p=0.05),
    transforms.ToTensor(),
    transforms.Normalize(MEAN, STD),
])

val_transforms = transforms.Compose([
    transforms.Resize((260, 260)),
    transforms.ToTensor(),
    transforms.Normalize(MEAN, STD),
])

# ─── DATASETS & LOADERS ───────────────────────────────────────────
train_ds = ImageFolder(os.path.join(CONFIG['data_dir'], 'train'), train_transforms)
val_ds   = ImageFolder(os.path.join(CONFIG['data_dir'], 'val'),   val_transforms)

train_loader = DataLoader(train_ds, batch_size=CONFIG['batch_size'],
                          shuffle=True, num_workers=4, pin_memory=True)
val_loader   = DataLoader(val_ds,   batch_size=CONFIG['batch_size'],
                          shuffle=False, num_workers=4, pin_memory=True)

print(f"Training samples:   {len(train_ds)}")
print(f"Validation samples: {len(val_ds)}")
print(f"Classes: {train_ds.classes}")

# ─── MODEL SETUP ──────────────────────────────────────────────────
def build_model(num_classes, freeze_backbone=True):
    """Build EfficientNet-B2 with custom head"""

    model = models.efficientnet_b2(weights='IMAGENET1K_V1')

    if freeze_backbone:
        for param in model.parameters():
            param.requires_grad = False

    # Replace classifier
    in_features = model.classifier[1].in_features  # 1408 for B2

    model.classifier = nn.Sequential(
        nn.Dropout(p=0.3),
        nn.Linear(in_features, 512),
        nn.BatchNorm1d(512),
        nn.ReLU(),
        nn.Dropout(p=0.2),
        nn.Linear(512, num_classes)
    )

    return model


# ─── TRAINING UTILITIES ───────────────────────────────────────────
def run_epoch(model, loader, optimizer, criterion, device, is_train=True):
    model.train() if is_train else model.eval()

    total_loss, correct, total = 0.0, 0, 0

    context = torch.enable_grad() if is_train else torch.no_grad()

    with context:
        for images, labels in loader:
            images, labels = images.to(device), labels.to(device)

            outputs = model(images)
            loss    = criterion(outputs, labels)

            if is_train:
                optimizer.zero_grad()
                loss.backward()
                optimizer.step()

            total_loss += loss.item()
            preds       = outputs.argmax(dim=1)
            correct    += (preds == labels).sum().item()
            total      += labels.size(0)

    return total_loss / len(loader), 100.0 * correct / total


# ─── PHASE 1: FEATURE EXTRACTION ──────────────────────────────────
print("\n" + "="*55)
print("  PHASE 1: Feature Extraction (Warming Up Head)")
print("="*55)

model = build_model(CONFIG['num_classes'], freeze_backbone=True)
model = model.to(CONFIG['device'])

optimizer1 = optim.Adam(
    filter(lambda p: p.requires_grad, model.parameters()),
    lr=CONFIG['phase1_lr']
)
criterion  = nn.CrossEntropyLoss(label_smoothing=0.1)
scheduler1 = optim.lr_scheduler.OneCycleLR(
    optimizer1, max_lr=CONFIG['phase1_lr'],
    steps_per_epoch=len(train_loader),
    epochs=CONFIG['phase1_epochs']
)

best_val_acc = 0.0

for epoch in range(1, CONFIG['phase1_epochs'] + 1):
    t0 = time.time()
    tr_loss, tr_acc = run_epoch(model, train_loader, optimizer1, criterion, CONFIG['device'])
    vl_loss, vl_acc = run_epoch(model, val_loader, None, criterion, CONFIG['device'], False)
    scheduler1.step()
    elapsed = time.time() - t0

    print(f"Ep {epoch:02d}/{CONFIG['phase1_epochs']} | "
          f"Train {tr_acc:.1f}% | Val {vl_acc:.1f}% | {elapsed:.1f}s")

    if vl_acc > best_val_acc:
        best_val_acc = vl_acc
        torch.save(model.state_dict(), CONFIG['model_save'])
        print(f"   ✅ Saved! Best val: {best_val_acc:.1f}%")


# ─── PHASE 2: FINE-TUNING ─────────────────────────────────────────
print("\n" + "="*55)
print("  PHASE 2: Fine-Tuning (Adapting Backbone)")
print("="*55)

# Load best weights from phase 1
model.load_state_dict(torch.load(CONFIG['model_save']))

# Unfreeze last 2 feature blocks + classifier
for param in model.parameters():
    param.requires_grad = False  # Re-freeze all

# Selectively unfreeze deep blocks
blocks = list(model.features.children())
for block in blocks[6:]:          # Unfreeze last 2 blocks (6, 7)
    for param in block.parameters():
        param.requires_grad = True

for param in model.classifier.parameters():
    param.requires_grad = True

trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
print(f"Fine-tuning {trainable:,} parameters")

optimizer2 = optim.AdamW([
    {'params': model.features[6:].parameters(),
     'lr': CONFIG['phase2_lr_backbone']},
    {'params': model.classifier.parameters(),
     'lr': CONFIG['phase2_lr_head']},
], weight_decay=CONFIG['weight_decay'])

scheduler2 = optim.lr_scheduler.CosineAnnealingLR(
    optimizer2, T_max=CONFIG['phase2_epochs'], eta_min=1e-7
)

for epoch in range(1, CONFIG['phase2_epochs'] + 1):
    t0 = time.time()
    tr_loss, tr_acc = run_epoch(model, train_loader, optimizer2, criterion, CONFIG['device'])
    vl_loss, vl_acc = run_epoch(model, val_loader, None, criterion, CONFIG['device'], False)
    scheduler2.step()
    elapsed = time.time() - t0

    print(f"Ep {epoch:02d}/{CONFIG['phase2_epochs']} | "
          f"Train {tr_acc:.1f}% | Val {vl_acc:.1f}% | {elapsed:.1f}s")

    if vl_acc > best_val_acc:
        best_val_acc = vl_acc
        torch.save(model.state_dict(), CONFIG['model_save'])
        print(f"   🏆 New Best! Val: {best_val_acc:.1f}%")

print(f"\n✅ Training complete! Best Val Accuracy: {best_val_acc:.1f}%")

Prediction Script (predict.py)

"""
predict.py — Load fine-tuned model and classify a dog image.
"""

import torch
import torchvision.models as models
import torchvision.transforms as transforms
from PIL import Image
import torch.nn as nn

CLASS_NAMES = [
    'beagle', 'bulldog', 'chihuahua', 'doberman', 'golden_retriever',
    'husky', 'labrador', 'poodle', 'rottweiler', 'yorkie'
]

MEAN = [0.485, 0.456, 0.406]
STD  = [0.229, 0.224, 0.225]

def load_model(weights_path, num_classes=10):
    model = models.efficientnet_b2(weights=None)  # No pretrained this time
    in_features = model.classifier[1].in_features

    model.classifier = nn.Sequential(
        nn.Dropout(p=0.3),
        nn.Linear(in_features, 512),
        nn.BatchNorm1d(512),
        nn.ReLU(),
        nn.Dropout(p=0.2),
        nn.Linear(512, num_classes)
    )

    model.load_state_dict(torch.load(weights_path, map_location='cpu'))
    model.eval()
    return model


def predict(image_path, model):
    transform = transforms.Compose([
        transforms.Resize((260, 260)),
        transforms.ToTensor(),
        transforms.Normalize(MEAN, STD),
    ])

    image  = Image.open(image_path).convert('RGB')
    tensor = transform(image).unsqueeze(0)   # Add batch dimension

    with torch.no_grad():
        outputs = model(tensor)
        probs   = torch.softmax(outputs, dim=1)
        top5    = torch.topk(probs, 5)

    print(f"\n🐕 Predictions for: {image_path}")
    print("-" * 35)

    for prob, idx in zip(top5.values[0], top5.indices[0]):
        bar = "█" * int(prob.item() * 30)
        print(f"  {CLASS_NAMES[idx]:20s} {bar} {prob.item()*100:.1f}%")

    predicted_class = CLASS_NAMES[top5.indices[0][0].item()]
    confidence      = top5.values[0][0].item() * 100
    print(f"\n✅ Result: {predicted_class} ({confidence:.1f}% confidence)")
    return predicted_class


# Run prediction
model  = load_model('best_dog_classifier.pth')
result = predict('test_dog.jpg', model)

Sample Output:

🐕 Predictions for: test_dog.jpg
-----------------------------------
  labrador             ██████████████████████ 76.3%
  golden_retriever     █████ 15.2%
  beagle               ██ 5.1%
  bulldog              █ 2.1%
  husky                  1.3%

✅ Result: labrador (76.3% confidence)

🚫 Section 11: Common Mistakes & How to Avoid Them

❌ Mistake 1: Using Wrong Normalization Values

Every pretrained model expects input normalized with the same mean/std used during pretraining. For ImageNet models, always use:
mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]

Using random normalization values will give terrible results — the model "sees" completely wrong colors and scales.
❌ Mistake 2: Fine-Tuning with Too High a Learning Rate

Never use a learning rate of 1e-3 or higher when fine-tuning pretrained layers. This destroys the carefully learned weights in just a few batches. Rule of thumb: Phase 2 LR should be 10x–100x smaller than Phase 1 LR.
❌ Mistake 3: Skipping the Warmup Phase

Jumping straight to fine-tuning without first warming up the new head causes large random gradients to backpropagate through the pretrained layers, instantly wrecking them. Always warm up the head first (Phase 1)!
❌ Mistake 4: Forgetting to Set model.eval() During Inference

BatchNorm and Dropout behave differently in training vs evaluation mode. Always call model.eval() before inference and model.train() before training. Missing this causes inconsistent and wrong predictions.
❌ Mistake 5: Using a Massive Model for Tiny Data

Fine-tuning ViT-Large (307M params) on 500 images = massive overfitting. Match your model size to your data size. For <2000 samples: use EfficientNet-B0 or ResNet18. Larger models need more data to benefit from fine-tuning.
✅ Best Practices Summary:
  • ✅ Always use the correct preprocessing (normalization) for your pretrained model
  • ✅ Warm up the head first (Phase 1) before fine-tuning the backbone (Phase 2)
  • ✅ Use differential learning rates — lower for earlier layers
  • ✅ Use label smoothing (0.05–0.1) to prevent overconfident predictions
  • ✅ Use data augmentation generously — especially when data is small
  • ✅ Monitor both training and validation loss — catch overfitting early
  • ✅ Save the best checkpoint (val acc), not the last one
  • ✅ Use cosine LR scheduling for fine-tuning
  • ✅ Try timm library for the easiest pretrained model API

🏆 Hero-Level Summary: Everything You've Mastered!

  • 🔹 Transfer Learning = Borrow knowledge from massive pretrained models. Stand on giants' shoulders.
  • 🔹 Feature Extraction = Freeze all pretrained weights. Train only a small new head. Fast, works with tiny data.
  • 🔹 Fine-Tuning = Unfreeze some/all layers. Let the model adapt. Needs more data. Gets better accuracy.
  • 🔹 Two-Phase Strategy = Warm up head first → then fine-tune backbone. Industry gold standard.
  • 🔹 Differential LR = Different learning rates for different layer depths. Lower LR for early layers.
  • 🔹 Progressive Unfreezing = Unfreeze layer by layer, top to bottom. Gentlest and most effective.
  • 🔹 Top Vision Models = EfficientNet, ConvNeXt V2, ViT, DINOv2, CLIP.
  • 🔹 LoRA / QLoRA = Fine-tune 7B LLMs on a single GPU by training tiny adapter matrices.
  • 🔹 PEFT = Family of efficient fine-tuning methods. LoRA is the most popular member.
  • 🔹 Always normalize correctly! Wrong normalization ruins everything.
Remember: NO production AI system is trained from scratch. Every real-world AI engineer uses pretrained models as their foundation. You now know exactly how to use them — from simple feature extraction all the way to LoRA fine-tuning of billion-parameter LLMs ⚡🧠

Keep experimenting, keep building, and keep learning! 🐼✨

Comments