Imagine you want to teach a world-class surgeon a new procedure. You don't put them back through 12 years of medical school. You just teach them the new part.
That's the core idea behind LoRA and QLoRA — two techniques that let you fine-tune billion-parameter language models by updating only a tiny fraction of their weights, using hardware that fits on a single consumer GPU.
The Problem — Why Fine-Tuning Used to Be Impossible
Modern LLMs like LLaMA 3, Mistral, and Gemma have billions of parameters.
A parameter is simply a number the model learned during training —
and there are billions of them packed together to represent everything the model knows.
Traditional fine-tuning means updating every single parameter using gradient descent on your new training data. This sounds straightforward, but the hardware requirements are staggering:
- LLaMA 3 8B model → ~16 GB just to store the weights in float16. Full fine-tuning needs 4–6× that in GPU memory for gradients and optimiser states. That's roughly 80–100 GB of GPU memory — six A100s minimum.
- LLaMA 3 70B model → Full fine-tuning requires clusters of high-end GPUs that cost thousands of dollars per hour to rent. Completely out of reach for individuals, startups, or research teams.
- Training time → Even with the hardware, full fine-tuning on a large dataset can take days to weeks per training run.
The ML community needed a smarter approach. Something that could adapt a model to a new task without retraining all of its billions of weights from scratch.
The Insight Behind LoRA — Adaptation Is Low-Rank
Here is the key insight that makes LoRA work: when you fine-tune a model, most of the weight changes you make are redundant.
In mathematics, a matrix has a property called its rank. The rank tells you how much truly independent information the matrix contains. A full-rank 1000×1000 matrix has 1,000,000 independent values. But a rank-4 matrix has only 4 truly independent dimensions of variation — even if it looks like a million numbers on the surface.
💡 Think of it like this: You want to describe where every person in a city lives. You could list all 5 million addresses individually (high rank, very redundant). Or you could describe the city's street grid — just a handful of axes — and derive every address from that. Low-rank = capturing the essential structure without the redundancy. 🗺️
Research into fine-tuning showed that the weight changes needed to adapt a pre-trained model to a new task are inherently low-rank. You don't need to update every element of every weight matrix. You just need to capture the low-rank "direction" of the update.
What Is LoRA? — The Core Idea
LoRA (Low-Rank Adaptation) was introduced in the 2021 paper "LoRA: Low-Rank Adaptation of Large Language Models" by Hu et al. at Microsoft.
Instead of modifying the original model weights directly, LoRA adds a small pair of adapter matrices alongside each weight matrix it wants to adapt. These adapter matrices are tiny — and only they are trained. The original model weights are completely frozen.
📐 How LoRA Works — The Weight Update
(Original Weight)
d × d = FROZEN ❄️
(LoRA Adapter)
d×r + r×d = TRAINED ✅
(Effective Weight)
Full model behaviour 🎯
r = rank (typically 4, 8, 16, or 64). The smaller r is, the fewer parameters are trained. With r=8 and d=4096, LoRA trains 65,536 params instead of 16,777,216. That's a 256× reduction!
The LoRA Math — Made Simple
In a standard neural network layer, the output h is computed as:
h = W · x
Where W is a large weight matrix (e.g., 4096 × 4096 in a big LLM)
and x is the input.
With LoRA, instead of updating W directly, we keep W frozen
and add a trainable low-rank correction:
h = W · x + (B · A) · x
Where:
Ais a matrix of shaped × r(initialized randomly)Bis a matrix of shaper × d(initialized to all zeros)ris the rank — a small number you choose (4, 8, 16, 64...)Bstarts at zero so that at training start, the adapter adds nothing to the model — training begins from the original behaviour
Only A and B are trained.
After training, you can optionally merge them back into W:
W_new = W + B·A.
This means zero inference overhead — the merged model runs at exactly the same speed as the original.
LoRA — Key Hyperparameters You Must Understand
1. Rank (r)
The rank r controls how many parameters LoRA trains.
Lower rank = fewer parameters = faster training but less expressive adaptation.
Higher rank = more parameters = slower but can capture more complex adaptations.
- r = 4 → Very lightweight. Good for simple task adaptation on small datasets.
- r = 8 → A solid default for most fine-tuning tasks. Start here.
- r = 16 or 32 → Better for complex tasks requiring significant behaviour change.
- r = 64+ → Rarely needed. Often overfits on small datasets.
2. Alpha (lora_alpha)
Alpha scales the LoRA update. The effective scaling factor is alpha / r.
Setting alpha = 2 * r (e.g., r=8, alpha=16) is the most common convention.
Think of alpha as controlling how boldly the adapter updates the model.
lora_alpha = 2 × lora_r.
So if r=8, set alpha=16.
If r=16, set alpha=32.
This scaling keeps gradient magnitudes stable as you change rank.
3. Target Modules
LoRA doesn't have to be applied to every layer. You choose which weight matrices to adapt. In transformer models, the most impactful targets are:
- q_proj, k_proj, v_proj → Query, Key, Value projections in the attention layers
- o_proj → Output projection in attention
- gate_proj, up_proj, down_proj → Feed-forward network layers (MLP)
Applying LoRA to all of these targets typically gives the best adaptation quality.
Applying it only to q_proj and v_proj is the minimal configuration.
4. Dropout
lora_dropout applies dropout regularisation inside the LoRA adapter.
This prevents overfitting on small datasets.
Typical values: 0.05 to 0.1. Set to 0 if your dataset is large (50,000+ samples).
LoRA in Code — Your First Fine-Tuning Job
Step 1 — Install the Required Libraries
pip install transformers peft accelerate datasets bitsandbytes trl
These five libraries are the backbone of the modern open-source fine-tuning ecosystem:
- transformers → Load pre-trained LLMs from Hugging Face Hub
- peft → Parameter-Efficient Fine-Tuning — implements LoRA, QLoRA, and other adapters
- accelerate → Handles multi-GPU and mixed-precision training automatically
- datasets → Load and process text datasets
- bitsandbytes → Enables 4-bit and 8-bit quantization (required for QLoRA)
- trl → Transformer Reinforcement Learning — provides SFTTrainer for clean supervised fine-tuning
Step 2 — Load the Base Model with LoRA Config
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, get_peft_model, TaskType
# Load the base model (using a small model for this example)
model_name = "mistralai/Mistral-7B-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token # required for batch training
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto", # loads in float16 or bfloat16 automatically
device_map="auto" # distributes across available GPUs/CPU
)
print(f"Model parameters: {model.num_parameters():,}")
Output:
Model parameters: 7,241,732,096
7.2 billion parameters — all frozen after we apply LoRA! ❄️
Step 3 — Define the LoRA Configuration
from peft import LoraConfig, get_peft_model, TaskType
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM, # we are fine-tuning a causal language model
r=8, # rank — start here for most tasks
lora_alpha=16, # scaling = alpha / r = 2.0
lora_dropout=0.05, # light regularisation
# Apply LoRA to all major transformer weight matrices
target_modules=[
"q_proj",
"k_proj",
"v_proj",
"o_proj",
"gate_proj",
"up_proj",
"down_proj"
],
bias="none" # don't train bias terms (saves memory)
)
# Wrap the base model with LoRA adapters
model = get_peft_model(model, lora_config)
# Check how many parameters we're actually training
model.print_trainable_parameters()
Output:
trainable params: 41,943,040 || all params: 7,283,675,136 || trainable%: 0.5757
We are training only 0.58% of the model's parameters! 42 million instead of 7.2 billion. The other 7.16 billion weights remain completely frozen. 🎯
Step 4 — Prepare Your Training Dataset
from datasets import load_dataset
# Load a sample instruction dataset
dataset = load_dataset("tatsu-lab/alpaca", split="train[:2000]") # use first 2000 samples
def format_instruction(sample):
"""
Format each sample into the Alpaca instruction template.
This is the exact text the model will train on.
"""
if sample["input"]:
return (
f"### Instruction:\n{sample['instruction']}\n\n"
f"### Input:\n{sample['input']}\n\n"
f"### Response:\n{sample['output']}"
)
else:
return (
f"### Instruction:\n{sample['instruction']}\n\n"
f"### Response:\n{sample['output']}"
)
# Apply formatting and keep only the text column
dataset = dataset.map(lambda x: {"text": format_instruction(x)})
print(dataset[0]["text"])
Output:
### Instruction:
Give three tips for staying healthy.
### Response:
1. Eat a balanced diet with plenty of fruits, vegetables, and whole grains.
2. Exercise regularly — aim for at least 30 minutes of moderate activity most days.
3. Prioritise sleep; most adults need 7–9 hours per night for optimal health.
Step 5 — Train with SFTTrainer
from trl import SFTTrainer
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./lora-mistral-alpaca",
num_train_epochs=3,
per_device_train_batch_size=4, # adjust based on your GPU memory
gradient_accumulation_steps=4, # effective batch size = 4 × 4 = 16
learning_rate=2e-4, # slightly higher than full fine-tuning is fine with LoRA
warmup_ratio=0.03,
lr_scheduler_type="cosine",
logging_steps=10,
save_strategy="epoch",
bf16=True, # bfloat16 training (A100/H100) or set fp16=True for older GPUs
optim="adamw_torch_fused", # fastest optimiser for modern PyTorch
report_to="none" # set to "wandb" to enable experiment tracking
)
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
args=training_args,
train_dataset=dataset,
dataset_text_field="text",
max_seq_length=2048,
packing=False # set True for faster training on short samples
)
trainer.train()
Output (sample):
{'loss': 2.4821, 'learning_rate': 0.0002, 'epoch': 0.1}
{'loss': 1.8342, 'learning_rate': 0.00018, 'epoch': 0.5}
{'loss': 1.4109, 'learning_rate': 0.00012, 'epoch': 1.0}
{'loss': 1.1203, 'learning_rate': 0.00006, 'epoch': 2.0}
{'loss': 0.9821, 'learning_rate': 0.00001, 'epoch': 3.0}
Training complete! ✅
Loss dropping consistently — the model is learning! 📉
Step 6 — Save and Merge the Adapter
# Option A: Save just the LoRA adapter weights (tiny — only ~40MB!)
model.save_pretrained("./lora-adapter-only")
tokenizer.save_pretrained("./lora-adapter-only")
# Option B: Merge LoRA weights back into the base model (zero inference overhead)
from peft import PeftModel
# Reload the original base model
base_model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
# Load and merge the LoRA adapter into it
merged_model = PeftModel.from_pretrained(base_model, "./lora-adapter-only")
merged_model = merged_model.merge_and_unload() # merges B·A into W, removes adapter
# Save the merged full model
merged_model.save_pretrained("./lora-merged-model")
tokenizer.save_pretrained("./lora-merged-model")
print("Model saved! No inference overhead — runs at full base model speed. ⚡")
Output:
Model saved! No inference overhead — runs at full base model speed. ⚡
The Problem LoRA Didn't Fully Solve — Memory at Load Time
LoRA dramatically reduces the number of trainable parameters. But there's a subtlety that catches many beginners off guard:
You still need to load the entire base model into GPU memory to train. On a 7B parameter model in float16, that's ~14 GB of GPU memory — just for the weights. Add gradients for the LoRA parameters and the optimiser state, and you're looking at ~18–20 GB total.
This rules out consumer GPUs like the RTX 3090 (24 GB — barely fits), RTX 3080 (10 GB — too small), or anything older.
Fine-tuning a 13B or 70B model is still far out of reach. The ML community needed one more breakthrough. That breakthrough was QLoRA.
What Is QLoRA? — LoRA Meets Quantization
QLoRA (Quantization-aware Low-Rank Adaptation) was introduced in the 2023 paper "QLoRA: Efficient Finetuning of Quantized LLMs" by Dettmers et al. at the University of Washington.
QLoRA combines two powerful ideas:
- 4-bit Quantization → Store the base model weights in 4 bits instead of 16 bits. This reduces the base model memory footprint by 4× with minimal quality loss.
- LoRA on top of the quantized model → Train LoRA adapters on top of the 4-bit frozen base model. The adapters themselves are stored in higher precision (bf16).
The result is extraordinary: you can fine-tune a 65B parameter model on a single 48 GB GPU — a task that previously required over 780 GB of GPU memory.
📊 Memory Comparison — Full FT vs LoRA vs QLoRA (7B Model)
Full Fine-Tuning
~80 GB
Weights + Gradients
+ Optimiser states
Needs 4× A100s minimum
LoRA (fp16)
~18 GB
fp16 weights + LoRA params
+ Optimiser state
Needs 1× A100 (40GB)
QLoRA (4-bit)
~6 GB
4-bit weights + bf16 LoRA
+ Optimiser state
Fits on RTX 3060 (12GB) ✅
The Three Innovations Inside QLoRA
QLoRA isn't just "quantize, then add LoRA." It introduced three specific innovations that make it work in practice:
Innovation 1 — NF4 (NormalFloat4) Quantization
Standard 4-bit quantization loses significant precision. QLoRA introduced a new data type called NF4 (NormalFloat4).
NF4 is specifically designed for neural network weights, which are typically normally distributed (bell-curve shaped) around zero. By placing more of the 4-bit number space around the centre of the distribution (where most weights live), NF4 preserves more information than generic 4-bit formats.
💡 Analogy: Imagine you have a ruler that can only mark 16 tick marks (4 bits = 16 possible values). If most of your measurements fall between 0.1 and 0.3, a smart ruler would crowd its tick marks in that range instead of spacing them evenly from 0 to 1. That's exactly what NF4 does for neural network weights. 📏
Innovation 2 — Double Quantization
When you quantize a tensor, you also need to store "quantization constants" (scaling factors that allow you to convert back to float). These constants themselves take memory.
QLoRA applies a second round of quantization to these quantization constants. This "double quantization" saves an additional ~0.5 GB per 7B model — small on its own, but significant when you're tight on VRAM.
Innovation 3 — Paged Optimisers
Training occasionally encounters memory spikes — moments where the GPU briefly needs more memory than usual (e.g., during a long sequence or a large gradient update).
QLoRA uses NVIDIA's unified memory to handle these spikes gracefully. When the GPU runs out of memory, it automatically "pages" some data to CPU RAM and pages it back when needed — exactly like how an operating system uses swap space. This prevents out-of-memory crashes during training.
QLoRA in Code — Fine-Tuning a 7B Model on a Single Consumer GPU
Step 1 — Load the Model in 4-bit (NF4)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, TaskType, prepare_model_for_kbit_training
model_name = "mistralai/Mistral-7B-v0.1"
# Configure 4-bit NF4 quantization
bnb_config = BitsAndBytesConfig(
load_in_4bit=True, # load weights in 4-bit (NF4)
bnb_4bit_quant_type="nf4", # use NF4 — the QLoRA data type
bnb_4bit_compute_dtype=torch.bfloat16, # compute in bf16 for speed
bnb_4bit_use_double_quant=True # enable double quantization to save more memory
)
# Load the model in 4-bit
model = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=bnb_config,
device_map="auto"
)
# Load the tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right" # important for causal LM training
print(f"Model loaded in 4-bit. Memory footprint: {model.get_memory_footprint() / 1e9:.2f} GB")
Output:
Model loaded in 4-bit. Memory footprint: 4.07 GB
4 GB for a 7 billion parameter model! 🤯 The same model in float16 would be 14 GB. We saved 10 GB through quantization alone.
Step 2 — Prepare the Model for k-bit Training
# This is a crucial step unique to QLoRA — it makes the quantized model
# trainable by casting certain layer norms to float32 for stability
model = prepare_model_for_kbit_training(model)
# Verify the model is ready
print("Gradient checkpointing enabled:", model.is_gradient_checkpointing)
print("Model dtype:", next(model.parameters()).dtype)
Output:
Gradient checkpointing enabled: True
Model dtype: torch.uint8
Step 3 — Add LoRA on Top of the Quantized Model
from peft import LoraConfig, get_peft_model
# Same LoRA config as before — but now applied on top of a 4-bit model
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=16, # slightly higher rank since the model is compressed
lora_alpha=32,
lora_dropout=0.05,
target_modules=[
"q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"
],
bias="none"
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
Output:
trainable params: 83,886,080 || all params: 3,827,765,248 || trainable%: 2.1913
Only 2.19% of parameters are trainable! The 4-bit frozen weights account for the vast majority — and they cost only a fraction of their normal memory. 🎯
Step 4 — Train with SFTTrainer (QLoRA-optimised)
from trl import SFTTrainer
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./qlora-mistral-output",
num_train_epochs=3,
per_device_train_batch_size=2, # smaller batch due to 4-bit constraints
gradient_accumulation_steps=8, # effective batch = 2 × 8 = 16
learning_rate=2e-4,
warmup_ratio=0.03,
lr_scheduler_type="cosine",
logging_steps=25,
save_strategy="epoch",
bf16=True,
# Key for QLoRA — use paged AdamW to handle memory spikes gracefully
optim="paged_adamw_32bit",
max_grad_norm=0.3, # gradient clipping for stability with 4-bit
report_to="none"
)
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
args=training_args,
train_dataset=dataset,
dataset_text_field="text",
max_seq_length=2048,
packing=False
)
trainer.train()
trainer.save_model("./qlora-final-adapter")
Output (sample):
{'loss': 2.3041, 'learning_rate': 0.0002, 'epoch': 0.2}
{'loss': 1.7823, 'learning_rate': 0.00015, 'epoch': 1.0}
{'loss': 1.2901, 'learning_rate': 0.00008, 'epoch': 2.0}
{'loss': 1.0312, 'learning_rate': 0.00001, 'epoch': 3.0}
Adapter saved to ./qlora-final-adapter ✅
paged_adamw_32bit uses NVIDIA's paged memory management to handle these spikes gracefully —
it's a key ingredient of QLoRA and should always be used when training in 4-bit.
Running Inference After QLoRA Fine-Tuning
Once trained, you can run inference in two ways — with the adapter still attached to the quantized model (minimal memory), or after merging into a dequantized model (for deployment).
Inference with the Adapter Attached
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, pipeline
import torch
# Reload base model in 4-bit
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
base_model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-v0.1",
quantization_config=bnb_config,
device_map="auto"
)
# Load the QLoRA adapter on top
model_with_adapter = PeftModel.from_pretrained(base_model, "./qlora-final-adapter")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1")
# Create a text generation pipeline
pipe = pipeline(
"text-generation",
model=model_with_adapter,
tokenizer=tokenizer,
max_new_tokens=256,
temperature=0.7,
do_sample=True
)
# Run inference
prompt = "### Instruction:\nExplain what gradient descent is in simple terms.\n\n### Response:\n"
result = pipe(prompt)[0]["generated_text"]
print(result)
Output:
### Instruction:
Explain what gradient descent is in simple terms.
### Response:
Gradient descent is like trying to find the lowest point in a hilly landscape while blindfolded.
At each step, you feel the ground around your feet to detect which direction slopes downward,
then take a small step in that direction. You keep repeating this — feel, step, feel, step —
until you stop going down and reach a valley.
In machine learning, the "landscape" is the loss function, the "valley" is the minimum loss,
and each "step" is an update to the model's weights based on the gradient (the slope of the loss).
Beautiful, clear explanation — generated by our fine-tuned model on a consumer GPU! 🎉
LoRA vs QLoRA — When to Use Which
Now that you understand both techniques, here's a practical decision guide for choosing between them:
-
Use LoRA when:
- You have access to a 24 GB+ GPU (e.g., RTX 3090, RTX 4090, A100)
- Your model is 7B or smaller and fits comfortably in fp16
- You want the fastest training speed without quantization overhead
- You plan to serve the merged model at low latency
-
Use QLoRA when:
- You have limited GPU memory (8–24 GB) — a standard consumer GPU
- You want to fine-tune a 13B, 34B, or 70B model on accessible hardware
- GPU rental cost is a concern and you want to minimise training time on expensive instances
- You are comfortable with a slight quality trade-off from 4-bit precision (usually negligible)
The Full Pipeline — LoRA/QLoRA in a Real MLOps System
🔄 LoRA/QLoRA in an MLOps Fine-Tuning Pipeline
1. Data Collection & Formatting
Curate instruction-response pairs → format as Alpaca or chat template → push to Hugging Face Dataset
2. Experiment Tracking (Weights & Biases / MLflow)
Log hyperparameters (r, alpha, lr), training curves, and adapter artifacts for every run
3. QLoRA Fine-Tuning (SFTTrainer)
Train adapter on quantized base model → save adapter weights (~40MB) after each epoch
4. Evaluation (ROUGE / BLEU / LLM-as-a-Judge)
Benchmark on held-out test set → compare against base model and previous adapter versions
5. Merge & Deploy
merge_and_unload() → push merged model to Hugging Face Hub → serve with vLLM or TGI
Experiment Tracking with Weights & Biases
import wandb
# Initialise a W&B run for this fine-tuning experiment
wandb.init(
project="qlora-fine-tuning",
name="mistral-7b-alpaca-r16",
config={
"model": "mistralai/Mistral-7B-v0.1",
"dataset": "alpaca-2000",
"lora_r": 16,
"lora_alpha": 32,
"learning_rate": 2e-4,
"epochs": 3,
"batch_size": 16,
"quantization": "4-bit NF4"
}
)
# Then in TrainingArguments, set:
# report_to="wandb"
# This automatically logs all metrics to your W&B dashboard
report_to="wandb" in TrainingArguments and it logs automatically.
Advanced Techniques — Taking LoRA Further
Technique 1 — LoRA+
LoRA+ (2024) is a simple improvement that sets different learning rates for the A and B matrices in each LoRA adapter. Research showed that setting the B matrix's learning rate 16× higher than the A matrix's consistently improves fine-tuning quality by 1–2% with zero additional cost.
from peft import LoraConfig
# Enable LoRA+ by setting the learning rate ratio
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
# LoRA+ setting: B matrix learns 16× faster than A matrix
loftq_config=None,
)
# Then in the optimiser, apply different learning rates
# (Using TRL's SFTTrainer, set loraplus_lr_ratio=16 in training config)
Technique 2 — RSLoRA (Rank-Stabilised LoRA)
Standard LoRA scales the adapter by alpha / r.
RSLoRA (2023) changes this to alpha / sqrt(r).
This seemingly small change means you can use higher ranks (r=64, r=128)
without the training becoming unstable.
Enabled in PEFT with a single flag:
lora_config = LoraConfig(
r=64, # can now safely use higher ranks
lora_alpha=64,
use_rslora=True, # enables rank-stabilised scaling: alpha / sqrt(r)
target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
task_type="CAUSAL_LM"
)
Technique 3 — DoRA (Weight-Decomposed Low-Rank Adaptation)
DoRA (2024) decomposes each weight matrix into its magnitude and direction components, then applies LoRA only to the direction. This makes the adapter learn more like full fine-tuning while keeping the parameter count low. Also a single flag in PEFT:
lora_config = LoraConfig(
r=16,
lora_alpha=32,
use_dora=True, # enables DoRA decomposition
target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
task_type="CAUSAL_LM"
)
use_rslora=True if using high ranks (r ≥ 32),
and try use_dora=True if you have the budget for a second training run comparison.
Both are single-line changes in PEFT that often yield measurable quality improvements.
Common Mistakes and How to Avoid Them 🪲
-
Forgetting
prepare_model_for_kbit_training()→ With QLoRA, skipping this step causes numerical instability and garbage outputs. It's mandatory — it enables gradient checkpointing and casts norms to float32. - Setting learning rate too high → LoRA is sensitive to learning rate. Values above 5e-4 often cause catastrophic forgetting — the model stops sounding coherent. Stick to 1e-4 to 3e-4 range as your starting point.
-
Using the wrong padding side →
For causal LM training (GPT-style models), always set
tokenizer.padding_side = "right". Left-padding causes the model to confuse the attention mask position during training and produces worse results. -
Training on too few steps without a warmup →
Without a warmup phase (even just 3% of total steps),
the first few batches can corrupt the adapter weights with extreme gradient updates.
Always set
warmup_ratio=0.03. -
Not using gradient clipping →
With 4-bit quantization, gradient spikes are more common.
Always set
max_grad_norm=0.3inTrainingArgumentsfor QLoRA runs. - Saving the full model instead of the adapter → The LoRA adapter file is typically 40–160 MB. The merged full model is 14 GB+. During experimentation, always save only the adapter and merge only when ready to deploy.
prepare_model_for_kbit_training() enables it automatically,
but verify with model.is_gradient_checkpointing == True before starting training.
Evaluating Your Fine-Tuned Model
Once training is done, how do you know if it actually improved? Here are three evaluation strategies from simple to comprehensive:
1. Qualitative Spot-Checking (Always Do This First)
test_prompts = [
"### Instruction:\nExplain the difference between RAM and storage in simple terms.\n\n### Response:\n",
"### Instruction:\nWrite a Python function that checks if a number is prime.\n\n### Response:\n",
"### Instruction:\nWhat are three benefits of drinking enough water daily?\n\n### Response:\n"
]
for prompt in test_prompts:
result = pipe(prompt, max_new_tokens=200)[0]["generated_text"]
# Extract only the response part
response = result.split("### Response:\n")[-1].strip()
print(f"PROMPT: {prompt.split(chr(10))[1]}")
print(f"RESPONSE: {response}\n{'-'*60}")
2. ROUGE Score for Summarisation Tasks
from evaluate import load
rouge = load("rouge")
# Generate responses for all test samples
predictions = []
references = []
for sample in test_dataset:
generated = pipe(sample["prompt"])[0]["generated_text"]
response = generated.split("### Response:\n")[-1].strip()
predictions.append(response)
references.append(sample["expected_output"])
results = rouge.compute(predictions=predictions, references=references)
print(f"ROUGE-1: {results['rouge1']:.4f}")
print(f"ROUGE-2: {results['rouge2']:.4f}")
print(f"ROUGE-L: {results['rougeL']:.4f}")
Output:
ROUGE-1: 0.4821
ROUGE-2: 0.2934
ROUGE-L: 0.4102
3. LLM-as-a-Judge for Open-Ended Quality
from openai import OpenAI
client = OpenAI()
def judge_response(instruction, response):
prompt = f"""Rate the following AI response to the given instruction.
Score from 1-5 where:
5 = Perfect: accurate, clear, complete, well-structured
4 = Good: mostly correct with minor gaps
3 = Average: partially helpful but missing key points
2 = Poor: significant errors or very incomplete
1 = Useless: wrong or irrelevant
Instruction: {instruction}
Response: {response}
Return ONLY JSON: {{"score": X, "reason": "brief explanation"}}"""
result = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
temperature=0
)
return result.choices[0].message.content
# Evaluate a sample
score_json = judge_response(
"Explain gradient descent to a 10-year-old.",
"Gradient descent is like rolling a ball downhill to find the lowest point..."
)
print(score_json)
Output:
{"score": 5, "reason": "Clear analogy appropriate for a child, accurately captures the optimization concept."}
Quick Summary 📝
What we covered today:
- The Problem → Full fine-tuning of billion-parameter LLMs requires 80–100+ GB GPU memory. Inaccessible for most teams and individuals.
- LoRA → Freezes the base model. Adds tiny trainable adapter matrices (B×A) alongside each weight matrix. Trains only 0.5–2% of total parameters. Zero inference overhead after merging.
- Key LoRA Hyperparameters → Rank (r), Alpha (lora_alpha = 2×r), target modules, dropout. Start with r=8, alpha=16.
- QLoRA → Loads the base model in 4-bit NF4. Applies LoRA on top. Enables fine-tuning of 7B models on 6 GB GPU, 70B models on 48 GB GPU.
- Three QLoRA Innovations → NF4 quantization, double quantization, paged optimisers.
- The PEFT + TRL Stack → LoraConfig + get_peft_model + SFTTrainer. The standard open-source fine-tuning pipeline.
- Advanced Variants → LoRA+, RSLoRA, DoRA — each a single flag change in PEFT, often improving quality at no extra cost.
- MLOps Integration → Track with W&B, save adapters to Hugging Face Hub, merge before deployment, serve with vLLM or TGI.
LoRA and QLoRA have democratised LLM fine-tuning. What once required a cluster of A100s and a six-figure compute budget now runs on hardware you might already have at home. That's not a small shift — it's a complete transformation in who gets to build and customise powerful AI systems 🧠✨
Comments
Post a Comment