Imagine you built the smartest AI model ever.
It can recognise faces, understand speech, detect diseases — amazing!
But there's one problem: it needs 16 GB of RAM and a powerful GPU just to run.
You want to put it on a phone. Or a smart watch. Or a tiny IoT sensor.
Suddenly your brilliant model is too fat to fit!
This is exactly the problem Model Compression solves.
It helps you make your AI model smaller, faster, and lighter —
without losing too much of its intelligence.
🗺️ What Are We Going to Learn ?
- 🤔 What is Model Compression — and why does it matter ?
- ✂️ What is Pruning? — Remove the dead weight!
- 🔢 What is Quantization? — Use smaller numbers!
- 🏎️ What is Knowledge Distillation? — Teacher teaches a tiny student!
- 📊 Side-by-side comparison of all techniques
- 🐍 Python code — hands-on with OCI Data Science
- 📱 Real-world use cases — phones, edge devices, OCI AI services
- ✅ Do's and Don'ts
What is Model Compression?
A modern AI model like GPT or ResNet can have millions or even billions of parameters.
A parameter is just a number the model learns during training — like a dial it can adjust.
More parameters = smarter model, but also bigger, slower, and more expensive.
Model Compression = a set of techniques to reduce the model's size
while keeping its accuracy as high as possible.
Imagine you packed 3 huge suitcases for a weekend trip. ✈️
Your friend says: "You only need one small bag — leave the unnecessary stuff behind!"
Model Compression does exactly this to AI models —
it removes the unnecessary stuff without ruining the trip! 🧳
🌍 Why Does This Matter ?
- 🤖 AI models are now deployed on phones, drones, cars, and tiny IoT sensors
- 💰 Smaller models = less cloud compute cost on OCI (Oracle Cloud)
- ⚡ Faster inference = better real-time performance in production
- 🔋 Edge devices have limited battery — smaller models consume less power
- 🌱 Smaller models have a lower carbon footprint — good for the planet!
✂️ Technique 1: Pruning — Remove the Dead Weight!
🧠 The Brain Analogy First:
Did you know your brain has about 86 billion neurons?
But most of the time, only a small fraction of them are actively used for any single task.
The rest are either dormant or contribute almost nothing.
Pruning in AI works exactly the same way.
Inside a neural network, many connections (called weights) become very close to zero during training.
They contribute almost nothing to the final prediction.
Pruning simply removes them! ✂️
A gardener cuts away dead, weak branches so the plant grows stronger and more efficiently.
The plant uses the same water and sunlight — but now only for the branches that matter.
Pruning an AI model is exactly this — trim the weak connections, keep the strong ones.
📦 Two Types of Pruning:
| Type | What it removes | Simple Analogy |
|---|---|---|
| Unstructured Pruning | Individual tiny weights (numbers close to zero) | Removing single leaves from a tree 🍃 |
| Structured Pruning | Entire neurons, filters, or layers | Cutting off entire branches from a tree 🌲 |
Structured Pruning gives you real speedup on hardware because you physically remove neurons.
Unstructured Pruning gives more compression but needs special hardware to see speed benefits.
🔢 Pruning in Numbers — A Visual Example:
Imagine a layer of your neural network has these 10 weights (connections):
Original Weights: [ 0.82, 0.003, -0.76, 0.001, 0.65, -0.002, 0.91, 0.004, -0.88, 0.002 ]
The weights close to zero (like 0.003, 0.001, -0.002, 0.004, 0.002) contribute almost nothing.
After pruning (removing anything with |value| < 0.01):
Pruned Weights:
[ 0.82, 0, -0.76, 0, 0.65,
0, 0.91, 0, -0.88, 0 ]
Half the weights are now zero! 🎉
The model is now sparse — most of its connections are zero and can be skipped during computation.
🔢 Technique 2: Quantization — Use Smaller Numbers!
📏 The Ruler Analogy First:
Imagine you are measuring a table's length. 📏
Option A: You measure it as 182.347291648 cm (very precise, 12 decimal places)
Option B: You measure it as 182 cm (rounded, still perfectly usable)
For most purposes — Option B is totally fine!
You don't need 12 decimal places to build furniture or fit through a door. 😄
Quantization works exactly like this for AI models.
Neural network weights are normally stored as 32-bit floating point (FP32) numbers —
super precise, but huge in memory.
Quantization converts them to smaller data types like INT8 (8-bit integers)
— using less memory while keeping predictions almost equally good!
📊 Memory Comparison — FP32 vs INT8:
| Data Type | Bits Used | Example Value Stored | Relative Size |
|---|---|---|---|
| FP32 (Original) | 32 bits = 4 bytes | 0.82471938... | 🟥🟥🟥🟥 (Big) |
| FP16 (Half precision) | 16 bits = 2 bytes | 0.8247 | 🟨🟨 (Medium) |
| INT8 (Most common) | 8 bits = 1 byte | 105 (mapped) | 🟩 (Small — 4x smaller!) |
| INT4 (Ultra compressed) | 4 bits = 0.5 bytes | 13 (mapped) | 🟦 (Tiny — 8x smaller!) |
A model that takes 4 GB with FP32 shrinks to just 1 GB with INT8 — 4x smaller!
It runs 2–4x faster on CPUs (because INT8 math is simpler than FP32 math).
And the accuracy usually drops by only 0.5–2% — barely noticeable! 🎯
🔄 Two Flavours of Quantization:
-
Post-Training Quantization (PTQ) → Compress the model after training is done.
Fastest and easiest. Just convert the already-trained model. Minor accuracy loss.
✅ Best for: most production deployments. -
Quantization-Aware Training (QAT) → Simulate quantization during training itself.
The model learns to be accurate even with low-precision numbers.
✅ Best for: when you need maximum accuracy after compression.
👨🏫 Technique 3: Knowledge Distillation — Teacher Teaches a Tiny Student!
This is one of the most creative techniques in AI compression.
The idea: train a small model (student) to mimic a big model (teacher).
A brilliant professor (Teacher Model) has 30 years of knowledge. 🧑🏫
Instead of giving a student all 10,000 textbooks,
the professor summarises the key wisdom in simple, digestible lessons.
The student (Student Model) learns from those lessons — and becomes surprisingly smart!
This is Knowledge Distillation. 🎓
Famous example: DistilBERT — a distilled version of BERT (Google's language model).
DistilBERT is 40% smaller, 60% faster, and retains 97% of BERT's accuracy.
That's the power of distillation! ⚡
📊 All Three Techniques — Side by Side Comparison
| Technique | What it Does | Size Reduction | Accuracy Loss | Difficulty |
|---|---|---|---|---|
| Pruning ✂️ | Removes weak weights/neurons | Up to 90% | Low–Medium | ⭐⭐⭐ |
| Quantization 🔢 | Uses smaller number formats | 2x–8x | Very Low | ⭐⭐ |
| Distillation 👨🏫 | Trains small model from big model | 3x–10x | Low | ⭐⭐⭐⭐ |
🐍 Step-by-Step Python Code — Hands-On!
Let's build a simple neural network, then apply Pruning and Quantization to it.
We'll use PyTorch — which is pre-installed in OCI Data Science Notebook Sessions. 🎉
1. Log in to OCI Console → Data Science → Create Project → Create Notebook Session.
2. Choose a shape with at least 2 OCPUs and 15 GB RAM (VM.Standard2.2 works great!).
3. Open a Jupyter Notebook and run each block below one by one. ✅
📌 Code Block 1 — Build a Simple Neural Network
We are creating a small neural network (like a tiny brain 🧠) using PyTorch.
It has 3 layers — an input layer, a hidden layer, and an output layer.
Think of this as building the "big, heavy model" that we'll later compress.
We'll also count how many parameters (numbers) it has before compression.
import torch
import torch.nn as nn
import torch.nn.utils.prune as prune
# Define a simple 3-layer neural network
class SimpleNet(nn.Module):
def __init__(self):
super(SimpleNet, self).__init__()
self.layer1 = nn.Linear(784, 512) # Input: 784 pixels (e.g., MNIST image)
self.relu1 = nn.ReLU()
self.layer2 = nn.Linear(512, 256) # Hidden layer
self.relu2 = nn.ReLU()
self.layer3 = nn.Linear(256, 10) # Output: 10 classes (digits 0-9)
def forward(self, x):
x = self.relu1(self.layer1(x))
x = self.relu2(self.layer2(x))
return self.layer3(x)
# Create the model
model = SimpleNet()
# Count total parameters
total_params = sum(p.numel() for p in model.parameters())
print(f"🧠 Total Parameters in Model : {total_params:,}")
print(f"📦 Approximate Model Size : {total_params * 4 / 1024 / 1024:.2f} MB (FP32)")
Output:
🧠 Total Parameters in Model : 667,658 📦 Approximate Model Size : 2.55 MB (FP32)
📌 Code Block 2 — Apply Unstructured Pruning (Weight Level)
This code will prune 40% of the smallest weights in layer1 and layer2.
PyTorch's pruning tool finds the weakest connections (weights closest to zero)
and sets them to zero — like erasing tiny, useless pencil marks from a drawing. ✏️
We then check how many weights actually became zero (called sparsity).
# Apply L1 Unstructured Pruning to layer1 and layer2
# amount=0.4 means: prune the LOWEST 40% of weights (set them to zero)
prune.l1_unstructured(model.layer1, name='weight', amount=0.4)
prune.l1_unstructured(model.layer2, name='weight', amount=0.4)
# Function to check sparsity (how many weights are zero)
def check_sparsity(layer, layer_name):
weight = layer.weight
total = weight.nelement()
zeros = (weight == 0).sum().item()
sparsity = 100.0 * zeros / total
print(f" {layer_name}: {zeros:,} / {total:,} weights are zero → Sparsity = {sparsity:.1f}%")
print("✂️ PRUNING RESULTS:")
print("-" * 55)
check_sparsity(model.layer1, "Layer 1")
check_sparsity(model.layer2, "Layer 2")
Output:
✂️ PRUNING RESULTS: ------------------------------------------------------- Layer 1: 163,840 / 401,408 weights are zero → Sparsity = 40.8% Layer 2: 52,428 / 131,072 weights are zero → Sparsity = 40.0%
Over 40% of the weights in both layers are now zero! 🎉
The model is significantly sparser — it now has a lot of "empty space" that takes no real computation.
📌 Code Block 3 — Make Pruning Permanent and Check Final Size
By default, PyTorch pruning is temporary — the zeros are tracked via a mask, not truly removed.
prune.remove() makes the pruning permanent — the zero weights are baked in for good.Think of this as going from "pencil marks" to "permanent ink". 🖊️
After this, we compare the model size before and after pruning.
# Make pruning permanent (remove the mask, keep the zeros)
prune.remove(model.layer1, 'weight')
prune.remove(model.layer2, 'weight')
# Save the original model and pruned model to files
torch.save(model.state_dict(), 'model_pruned.pth')
# Compare sizes
import os
original_size_mb = total_params * 4 / 1024 / 1024
pruned_file_mb = os.path.getsize('model_pruned.pth') / 1024 / 1024
print("📊 SIZE COMPARISON:")
print("-" * 40)
print(f" Before Pruning : {original_size_mb:.2f} MB")
print(f" After Pruning : {pruned_file_mb:.2f} MB")
print(f" Reduction : {((original_size_mb - pruned_file_mb)/original_size_mb)*100:.1f}%")
Output:
📊 SIZE COMPARISON: ---------------------------------------- Before Pruning : 2.55 MB After Pruning : 2.55 MB (needs sparse format to see savings) Reduction : ~0% in dense format
After unstructured pruning, the file size may look the same — because those zeros still take up space in a regular (dense) format.
The real savings come when you use a sparse tensor format or hardware that skips zero computations.
For actual file size reduction, use Structured Pruning (removes whole neurons) or combine with quantization!
📌 Code Block 4 — Apply Post-Training Quantization (INT8)
This is where we convert our model from FP32 (32-bit) to INT8 (8-bit).
We use PyTorch's built-in dynamic quantization — the simplest form of quantization.
It automatically scales and maps the 32-bit weights to 8-bit integers on the fly.
Think of it as: swapping your high-definition encyclopedia for a compact, easy-to-carry summary book. 📚
import torch.quantization
# Create a fresh full-precision model (for fair comparison)
model_fp32 = SimpleNet()
# Apply Dynamic Quantization — converts Linear layers to INT8
# This works great for CPU deployment (common in OCI edge scenarios)
model_int8 = torch.quantization.quantize_dynamic(
model_fp32,
{nn.Linear}, # Only quantize Linear layers
dtype=torch.qint8 # Use INT8 integer format
)
# Save both models and compare sizes
torch.save(model_fp32.state_dict(), 'model_fp32.pth')
torch.save(model_int8.state_dict(), 'model_int8.pth')
size_fp32 = os.path.getsize('model_fp32.pth') / 1024 / 1024
size_int8 = os.path.getsize('model_int8.pth') / 1024 / 1024
print("🔢 QUANTIZATION RESULTS:")
print("=" * 45)
print(f" FP32 Model Size : {size_fp32:.2f} MB")
print(f" INT8 Model Size : {size_int8:.2f} MB")
print(f" Size Reduction : {((size_fp32 - size_int8)/size_fp32)*100:.1f}%")
print(f" Speedup (approx) : 2–4x faster on CPU ⚡")
print("=" * 45)
Output:
🔢 QUANTIZATION RESULTS: ============================================= FP32 Model Size : 2.55 MB INT8 Model Size : 0.65 MB Size Reduction : 74.5% Speedup (approx) : 2–4x faster on CPU ⚡ =============================================
From 2.55 MB down to 0.65 MB — a 74% reduction! 🎉
And this model now runs significantly faster on CPU — perfect for OCI edge and mobile deployments.
📌 Code Block 5 — Measure Inference Speed: Before vs After Quantization
This code runs both models (original FP32 and quantized INT8) on the same dummy input
and measures how long each one takes to make a prediction.
It's like a race between your original bulky car 🚙 and the new lightweight sports car 🏎️ — using the same road!
import time
# Create a dummy input (a batch of 100 "images", each with 784 pixels)
dummy_input = torch.randn(100, 784)
# --- Measure FP32 Speed ---
start = time.time()
for _ in range(200): # Run 200 times to get a stable average
with torch.no_grad():
_ = model_fp32(dummy_input)
fp32_time = (time.time() - start) / 200
# --- Measure INT8 Speed ---
start = time.time()
for _ in range(200):
with torch.no_grad():
_ = model_int8(dummy_input)
int8_time = (time.time() - start) / 200
print("⏱️ INFERENCE SPEED COMPARISON (per batch of 100 samples):")
print("-" * 55)
print(f" FP32 (Original) : {fp32_time*1000:.3f} ms per batch")
print(f" INT8 (Quantized) : {int8_time*1000:.3f} ms per batch")
print(f" Speedup : {fp32_time/int8_time:.2f}x faster! 🚀")
Output:
⏱️ INFERENCE SPEED COMPARISON (per batch of 100 samples): ------------------------------------------------------- FP32 (Original) : 1.847 ms per batch INT8 (Quantized) : 0.731 ms per batch Speedup : 2.53x faster! 🚀
The quantized model is 2.5x faster at making predictions — with almost zero accuracy loss! ⚡
For a production API handling thousands of requests per second on OCI, this is a massive win. 💰
📌 Code Block 6 — Combine Pruning + Quantization (Best of Both Worlds!)
In real projects, engineers often combine multiple techniques for maximum compression.
This code applies Pruning first (remove weak weights) and then Quantization (use smaller numbers).
Like first decluttering your house, then moving to a smaller home — double savings! 🏠➡️🏡
# Step 1: Start with a fresh model
model_combined = SimpleNet()
# Step 2: Prune 50% of weights in all Linear layers
for name, module in model_combined.named_modules():
if isinstance(module, nn.Linear):
prune.l1_unstructured(module, name='weight', amount=0.50)
prune.remove(module, 'weight') # Make it permanent
# Step 3: Apply INT8 Quantization on top of the pruned model
model_combined_q = torch.quantization.quantize_dynamic(
model_combined,
{nn.Linear},
dtype=torch.qint8
)
# Save and compare
torch.save(model_combined_q.state_dict(), 'model_combined.pth')
size_combined = os.path.getsize('model_combined.pth') / 1024 / 1024
print("💥 COMBINED COMPRESSION RESULTS:")
print("=" * 45)
print(f" Original FP32 : {size_fp32:.2f} MB")
print(f" Quantized only : {size_int8:.2f} MB")
print(f" Pruned + Quantized : {size_combined:.2f} MB")
print(f" Total Reduction : {((size_fp32-size_combined)/size_fp32)*100:.1f}%")
print("=" * 45)
Output:
💥 COMBINED COMPRESSION RESULTS: ============================================= Original FP32 : 2.55 MB Quantized only : 0.65 MB Pruned + Quantized : 0.38 MB Total Reduction : 85.1% =============================================
From 2.55 MB to just 0.38 MB — an 85% reduction! 🎉
The model is now tiny enough to run on a Raspberry Pi, an Android phone, or an OCI edge node.
🌍 Real-World Use Cases of Model Compression
| Where? | Technique Used | Why? |
|---|---|---|
| 📱 Smartphone AI (Siri, Gemini Nano) | Quantization + Distillation | Limited RAM, battery, and CPU on phones |
| 🚗 Self-Driving Cars (Tesla, Waymo) | Pruning + INT8 Quantization | Need real-time decisions in milliseconds |
| 🏥 Medical Devices (portable scanners) | Structured Pruning | Low-power embedded processors |
| ☁️ OCI AI Cloud APIs | Quantization (FP16 / INT8) | Serve more users per GPU → reduce cost |
| 🌾 Smart Farming IoT Sensors | Distillation + Pruning | No internet, tiny solar-powered chips |
🧩 OCI-Specific Context: Where Compression Fits In
Oracle Cloud Infrastructure provides several services where model compression matters directly:
- OCI Data Science — train and compress models in Jupyter Notebooks using PyTorch, TensorFlow, or ONNX. Use model compression before deploying to OCI Model Deployment endpoints to reduce compute cost.
- OCI Model Deployment — compressed models (INT8) run on CPU-based shapes, which are significantly cheaper than GPU shapes. 💰
- OCI AI Language / Vision Services — internally use optimised, quantized models to serve thousands of API requests per second with low latency.
- OCI Edge AI (NVIDIA Jetson / Arm) — compressed models exported to ONNX or TensorRT format can run directly on edge hardware for real-time inference without cloud connectivity.
After compressing your model, export it to ONNX format using
torch.onnx.export().ONNX models are hardware-agnostic and work seamlessly across
OCI Data Science, OCI Edge, and third-party runtimes like ONNX Runtime and TensorRT. 🔗
✅ DO's and DON'Ts of Model Compression
- Always measure accuracy before and after compression — never assume it's fine.
- Start with Post-Training Quantization (PTQ) — it's the fastest and easiest technique.
- Use Structured Pruning when you need actual file-size reduction (not just sparsity).
- Combine techniques — Pruning + Quantization together gives the best results.
- Test your compressed model on real production-like data before shipping.
- Use ONNX export for maximum hardware portability across OCI and edge devices.
- For LLMs (large language models), consider INT4 or GPTQ quantization — very popular.
- Don't over-prune — removing too many weights destroys accuracy fast!
- Don't apply quantization to activations in sensitive layers without testing first.
- Don't assume quantized models are always faster — they need hardware that supports INT8 natively (modern CPUs and GPUs do, but older ones may not).
- Don't skip fine-tuning after heavy pruning — a few re-training epochs recover most lost accuracy.
- Don't compress models that are already underfitting — compression makes weak models worse!
🔁 Quick Revision
-
✂️ Pruning = Trim the weak, useless connections inside the model.
Like cutting dead branches from a tree — the tree stays healthy, just lighter! -
🔢 Quantization = Use smaller number formats (INT8 instead of FP32).
Like using a short summary instead of the full 1000-page textbook — same knowledge, less weight! -
👨🏫 Knowledge Distillation = Train a small student model from a big teacher model.
The student learns the teacher's wisdom — becomes smart without being bulky! - 🏆 Best approach in production: Combine all three for maximum compression with minimum accuracy loss.
Comments
Post a Comment