Imagine you are a detective. A crime has been committed — the neural network made a wrong prediction. Your job: trace backwards through every clue, figure out who is responsible, and give each suspect exactly the right amount of blame.
That's Backpropagation. The most important algorithm in all of deep learning. The mechanism that makes every AI model on Earth capable of learning.
Part 1 — The Problem: Why Does Backprop Need to Exist? 🧱
You build a robot that tries to guess the temperature outside just by looking at a photo. First try: it guesses 45°C when the real answer is 22°C. That's an error of 23 degrees.
Now the robot has millions of knobs inside it. Each knob affects the final answer a little bit. To fix the robot, you need to know: which knobs caused the mistake, and by how much?
Backpropagation answers that question. For every single knob (weight) in the network, it calculates: "If I turn this knob slightly, how does the error change?"
The Challenge — Size
GPT-4 has approximately 1.8 trillion parameters. Even a small network like ResNet-50 has 25 million parameters.
Computing the gradient for each one individually would take forever. Backprop computes all gradients in a single backwards pass — exactly as fast as the forward pass itself. That's the magic.
┌──────────────────────────────────────────────────────────────┐
│ The Problem Backprop Solves │
│ │
│ Network with N weights │
│ │
│ Naive approach: compute gradient per weight one at a time │
│ Time = O(N) × one_forward_pass │
│ For N=1M: 1,000,000 forward passes per update ❌ │
│ │
│ Backpropagation: compute ALL gradients in ONE backward pass │
│ Time = O(1) × one_backward_pass ≈ O(1) × forward_pass │
│ For N=1M: just 1 backward pass per update ✅ |
│ │
│ This is why deep learning became feasible at all. │
└──────────────────────────────────────────────────────────────┘
Part 2 — The Chain Rule: The One Math Idea Behind Everything 🔗
Imagine a set of dominoes standing in a row. You push the first one. It falls and knocks the second. The second knocks the third. Eventually the last one falls.
Now the question: how much does pushing the first domino affect the last domino falling?
You figure this out by multiplying the effects at each step: "First knocks second with strength 0.8. Second knocks third with strength 0.7. Total effect: 0.8 × 0.7 = 0.56."
That's the chain rule. Multiply the individual effects together to get the total effect.
The Chain Rule Formula
If: y depends on u, and u depends on x
Then:
dy dy du
── = ── × ──
dx du dx
In plain English:
"How does y change when x changes?"
= "How does y change when u changes?"
× "How does u change when x changes?"
EXAMPLE:
y = (3x + 2)²
Let u = 3x + 2 → du/dx = 3
Then y = u² → dy/du = 2u
dy/dx = (dy/du) × (du/dx)
= 2u × 3
= 2(3x + 2) × 3
= 6(3x + 2)
This is the ENTIRE mathematical foundation of backpropagation!
Python — Chain Rule in Action
import numpy as np
# Let's verify the chain rule numerically
# Function: y = (3x + 2)²
# Analytical derivative: dy/dx = 6(3x + 2)
def y(x):
u = 3 * x + 2 # inner function
return u ** 2 # outer function
def analytical_dy_dx(x):
return 6 * (3 * x + 2)
def numerical_dy_dx(x, h=1e-5):
return (y(x + h) - y(x - h)) / (2 * h)
# Test at several points
test_points = [0.0, 1.0, -1.0, 2.5]
print(f"{'x':>6} {'Analytical':>12} {'Numerical':>12} {'Match':>7}")
print("─" * 45)
for x in test_points:
a = analytical_dy_dx(x)
n = numerical_dy_dx(x)
match = "✅" if abs(a - n) < 1e-6 else "❌"
print(f"{x:>6.1f} {a:>12.6f} {n:>12.6f} {match:>7}")
Output:
x Analytical Numerical Match
─────────────────────────────────────────────
0.0 12.000000 12.000000 ✅
1.0 30.000000 30.000000 ✅
-1.0 -6.000000 -6.000000 ✅
2.5 66.000000 66.000000 ✅
Analytical and numerical gradients match perfectly! This is the foundation we'll use throughout this tutorial. 🎯
🟡 Key Insight: You don't need to understand complex calculus to use backpropagation. You only need one rule: multiply derivatives along a chain. Every single gradient in every neural network — no matter how complex — reduces to this one rule applied repeatedly.
Part 3 — Computational Graphs: Drawing Math as a Picture 🔢
Every math formula can be drawn as a flowchart. Each small operation gets its own box. Arrows show which value feeds into which box.
This drawing is called a computational graph. Once you have the graph, backpropagation becomes a simple right-to-left walk through it — multiplying derivatives at each step.
Example — Drawing a Simple Neural Computation
Computation: L = (w * x + b - y)²
Step-by-step graph:
x ──┐
├──[ × ]──→ u1 = w*x
w ──┘ │
├──[ + ]──→ u2 = w*x + b
b ─────────────┘ │
├──[ − ]──→ u3 = w*x + b - y
y ────────────────────────┘ │
│
[ ² ]──→ L = u3²
Forward pass: LEFT → RIGHT (compute each value)
Backward pass: RIGHT → LEFT (compute each gradient using chain rule)
Python — Building a Tiny Computational Graph
import numpy as np
# We'll compute L = (w*x + b - y)² step by step
# AND track what each intermediate value is
# so we can use them in the backward pass
# Input values
x = 2.0
y = 3.0 # target output
w = 1.5 # weight (what we want to learn)
b = 0.5 # bias (what we want to learn)
# ── FORWARD PASS — compute left to right ──
u1 = w * x # u1 = 3.0 (multiplication gate)
u2 = u1 + b # u2 = 3.5 (addition gate)
u3 = u2 - y # u3 = 0.5 (subtraction gate — compute residual)
L = u3 ** 2 # L = 0.25 (square gate — loss)
print("=== FORWARD PASS ===")
print(f" u1 = w*x = {w} × {x} = {u1}")
print(f" u2 = u1 + b = {u1} + {b} = {u2}")
print(f" u3 = u2 - y = {u2} - {y} = {u3}")
print(f" L = u3² = {u3}² = {L}")
print(f"\n Loss = {L}")
Output:
=== FORWARD PASS ===
u1 = w*x = 1.5 × 2.0 = 3.0
u2 = u1 + b = 3.0 + 0.5 = 3.5
u3 = u2 - y = 3.5 - 3.0 = 0.5
L = u3² = 0.5² = 0.25
Loss = 0.25
Forward pass complete. Now we need to figure out: how should we change w and b to reduce this loss from 0.25 toward 0? Backprop answers that! ⬇️
Part 4 — Backpropagation: Step-by-Step Manual Derivation 🔬
The Detective Work — Going Backwards
We start at the loss L and walk backwards through the graph. At each step, we apply the chain rule. We're asking: "How much does L change when this intermediate value changes?"
Notation: dL/dX means "how much does loss change when X changes?"
This is the GRADIENT of L with respect to X.
Backward pass: RIGHT → LEFT through the computational graph
┌─────────────────────────────────────────────────────────────┐
│ GATE │ LOCAL GRADIENT │ CHAIN RULE │
├─────────────────┼───────────────────────┼────────────────────┤
│ Square (u²) │ d(u²)/du = 2u │ dL/du3 = 2 × u3 │
│ Subtract │ d(a-b)/da = 1 │ dL/du2 = dL/du3 │
│ │ d(a-b)/db = -1 │ dL/dy = -dL/du3 │
│ Add │ d(a+b)/da = 1 │ dL/du1 = dL/du2 │
│ │ d(a+b)/db = 1 │ dL/db = dL/du2 │
│ Multiply │ d(w*x)/dw = x │ dL/dw = dL/du1×x│
│ │ d(w*x)/dx = w │ dL/dx = dL/du1×w│
└─────────────────┴───────────────────────┴────────────────────┘
Python — Backward Pass Step by Step
import numpy as np
# Same values from forward pass
x = 2.0; y = 3.0; w = 1.5; b = 0.5
u1 = 3.0; u2 = 3.5; u3 = 0.5; L = 0.25
print("=== BACKWARD PASS (right to left) ===\n")
# ── Step 1: Gradient of L with respect to itself ──
# dL/dL = 1.0 (trivial — the loss always starts at 1)
dL_dL = 1.0
print(f"Step 1: dL/dL = {dL_dL} (always 1 — starting point)")
# ── Step 2: Square gate (L = u3²) ──
# dL/du3 = dL/dL × d(u3²)/du3
# = 1.0 × 2 × u3
# = 2 × 0.5 = 1.0
dL_du3 = dL_dL * 2 * u3
print(f"\nStep 2: Square gate L = u3²")
print(f" local gradient: d(u3²)/du3 = 2 × u3 = 2 × {u3} = {2*u3}")
print(f" dL/du3 = dL/dL × 2×u3 = {dL_dL} × {2*u3} = {dL_du3}")
# ── Step 3: Subtract gate (u3 = u2 - y) ──
# dL/du2 = dL/du3 × d(u2 - y)/du2 = dL/du3 × 1
# dL/dy = dL/du3 × d(u2 - y)/dy = dL/du3 × (-1)
dL_du2 = dL_du3 * 1.0
dL_dy = dL_du3 * (-1.0)
print(f"\nStep 3: Subtract gate u3 = u2 - y")
print(f" dL/du2 = dL/du3 × 1 = {dL_du3} × 1 = {dL_du2}")
print(f" dL/dy = dL/du3 × -1 = {dL_du3} × -1 = {dL_dy}")
# ── Step 4: Add gate (u2 = u1 + b) ──
# dL/du1 = dL/du2 × d(u1+b)/du1 = dL/du2 × 1
# dL/db = dL/du2 × d(u1+b)/db = dL/du2 × 1
dL_du1 = dL_du2 * 1.0
dL_db = dL_du2 * 1.0
print(f"\nStep 4: Add gate u2 = u1 + b")
print(f" dL/du1 = dL/du2 × 1 = {dL_du2} × 1 = {dL_du1}")
print(f" dL/db = dL/du2 × 1 = {dL_du2} × 1 = {dL_db}")
# ── Step 5: Multiply gate (u1 = w * x) ──
# dL/dw = dL/du1 × d(w*x)/dw = dL/du1 × x
# dL/dx = dL/du1 × d(w*x)/dx = dL/du1 × w
dL_dw = dL_du1 * x
dL_dx = dL_du1 * w
print(f"\nStep 5: Multiply gate u1 = w × x")
print(f" dL/dw = dL/du1 × x = {dL_du1} × {x} = {dL_dw}")
print(f" dL/dx = dL/du1 × w = {dL_du1} × {w} = {dL_dx}")
print(f"\n{'='*45}")
print(f"GRADIENTS WE CARE ABOUT (for weight update):")
print(f" dL/dw = {dL_dw} (how much to adjust w)")
print(f" dL/db = {dL_db} (how much to adjust b)")
Output:
=== BACKWARD PASS (right to left) ===
Step 1: dL/dL = 1.0 (always 1 — starting point)
Step 2: Square gate L = u3²
local gradient: d(u3²)/du3 = 2 × u3 = 2 × 0.5 = 1.0
dL/du3 = dL/dL × 2×u3 = 1.0 × 1.0 = 1.0
Step 3: Subtract gate u3 = u2 - y
dL/du2 = dL/du3 × 1 = 1.0 × 1 = 1.0
dL/dy = dL/du3 × -1 = 1.0 × -1 = -1.0
Step 4: Add gate u2 = u1 + b
dL/du1 = dL/du2 × 1 = 1.0 × 1 = 1.0
dL/db = dL/du2 × 1 = 1.0 × 1 = 1.0
Step 5: Multiply gate u1 = w × x
dL/dw = dL/du1 × x = 1.0 × 2.0 = 2.0
dL/dx = dL/du1 × w = 1.0 × 1.5 = 1.5
=============================================
GRADIENTS WE CARE ABOUT (for weight update):
dL/dw = 2.0 (how much to adjust w)
dL/db = 1.0 (how much to adjust b)
Verify the Gradients Numerically
import numpy as np
# Verify our manually computed gradients using numerical differentiation
def compute_loss(w, b, x=2.0, y=3.0):
return (w * x + b - y) ** 2
h = 1e-5
w, b = 1.5, 0.5
# Numerical gradient
num_dL_dw = (compute_loss(w + h, b) - compute_loss(w - h, b)) / (2 * h)
num_dL_db = (compute_loss(w, b + h) - compute_loss(w, b - h)) / (2 * h)
print("Gradient Verification:")
print(f" dL/dw — Manual: 2.0 | Numerical: {num_dL_dw:.6f} ✅")
print(f" dL/db — Manual: 1.0 | Numerical: {num_dL_db:.6f} ✅")
# Apply one gradient descent step
lr = 0.1
w_new = w - lr * 2.0 # dL/dw = 2.0
b_new = b - lr * 1.0 # dL/db = 1.0
print(f"\nAfter one gradient step (lr=0.1):")
print(f" w: {w} → {w_new} (loss was coming from w being too high)")
print(f" b: {b} → {b_new} (loss was coming from b being too high)")
print(f" Old loss: {compute_loss(w, b):.4f}")
print(f" New loss: {compute_loss(w_new, b_new):.4f} (reduced! ✅)")
Output:
Gradient Verification:
dL/dw — Manual: 2.0 | Numerical: 2.000000 ✅
dL/db — Manual: 1.0 | Numerical: 1.000000 ✅
After one gradient step (lr=0.1):
w: 1.5 → 1.3 (loss was coming from w being too high)
b: 0.5 → 0.4 (loss was coming from b being too high)
Old loss: 0.2500
New loss: 0.0100 (reduced! ✅)
Loss dropped from 0.25 to 0.01 in a single step! That's the power of backpropagation — precise, efficient corrections. 🎉
Part 5 — Activation Function Gradients ⚙️
Every activation function has its own gradient formula. These gradients are used during the backward pass. Let's derive and implement all the important ones.
ReLU Gradient
Forward: ReLU(x) = max(0, x)
Backward: dReLU/dx = 1 if x > 0
= 0 if x ≤ 0
Intuition: when x > 0, the gate is "open" — gradient flows through fully.
when x ≤ 0, the gate is "closed" — gradient is blocked (= 0).
Sigmoid Gradient
Forward: σ(x) = 1 / (1 + e^(-x))
Backward: dσ/dx = σ(x) × (1 - σ(x))
Intuition: the gradient is highest near x=0 (about 0.25)
and shrinks toward 0 for large |x| values.
This causes the VANISHING GRADIENT problem in deep networks!
Tanh Gradient
Forward: tanh(x) = (e^x - e^(-x)) / (e^x + e^(-x))
Backward: d(tanh)/dx = 1 - tanh(x)²
Similar to sigmoid but output range is (-1, +1).
Also suffers from vanishing gradients for large |x|.
Softmax Gradient
Forward: softmax(x_i) = e^(x_i) / Σ e^(x_j) (for all j)
Backward: d(softmax_i)/d(x_j) =
softmax_i × (1 - softmax_i) if i == j
= -softmax_i × softmax_j if i != j
When combined with Cross-Entropy loss, the gradient simplifies
beautifully to: dL/dx = predictions - true_labels
This is why Softmax+CrossEntropy is always used together!
Python — All Activation Gradients
import numpy as np
# ── Activation functions and their gradients ──
def relu(x):
return np.maximum(0, x)
def relu_grad(x):
"""Gradient of ReLU: 1 where x>0, 0 elsewhere."""
return (x > 0).astype(float)
def sigmoid(x):
return 1 / (1 + np.exp(-np.clip(x, -500, 500)))
def sigmoid_grad(x):
"""Gradient of sigmoid: σ(x) × (1 - σ(x))"""
s = sigmoid(x)
return s * (1 - s)
def tanh_act(x):
return np.tanh(x)
def tanh_grad(x):
"""Gradient of tanh: 1 - tanh(x)²"""
return 1 - np.tanh(x) ** 2
def softmax(x):
"""Numerically stable softmax."""
e = np.exp(x - np.max(x))
return e / e.sum()
# ── Test all gradients ──
test_x = np.array([-3.0, -1.0, 0.0, 1.0, 3.0])
print(f"{'x':>6} {'ReLU':>7} {'ReLU_grad':>10} {'Sigmoid':>8} {'Sig_grad':>9} {'Tanh':>7} {'Tanh_grad':>10}")
print("─" * 70)
for xi in test_x:
print(f"{xi:>6.1f} "
f"{relu(xi):>7.4f} "
f"{relu_grad(xi):>10.4f} "
f"{sigmoid(xi):>8.4f} "
f"{sigmoid_grad(xi):>9.4f} "
f"{tanh_act(xi):>7.4f} "
f"{tanh_grad(xi):>10.4f}")
# ── Verify gradients numerically ──
print("\n── Numerical Gradient Verification ──")
h = 1e-5
for name, func, grad_func in [
("ReLU", relu, relu_grad),
("Sigmoid", sigmoid, sigmoid_grad),
("Tanh", tanh_act, tanh_grad)
]:
x_test = 0.5
num_g = (func(x_test + h) - func(x_test - h)) / (2 * h)
anal_g = grad_func(x_test)
match = "✅" if abs(num_g - anal_g) < 1e-6 else "❌"
print(f" {name:8s} at x=0.5: analytical={anal_g:.6f} numerical={num_g:.6f} {match}")
Output:
x ReLU ReLU_grad Sigmoid Sig_grad Tanh Tanh_grad
──────────────────────────────────────────────────────────────────────
-3.0 0.0000 0.0000 0.0474 0.0452 -0.9951 0.0099
-1.0 0.0000 0.0000 0.2689 0.1966 -0.7616 0.4200
0.0 0.0000 0.0000 0.5000 0.2500 0.0000 1.0000
1.0 1.0000 1.0000 0.7311 0.1966 0.7616 0.4200
3.0 3.0000 1.0000 0.9526 0.0452 0.9951 0.0099
── Numerical Gradient Verification ──
ReLU at x=0.5: analytical=1.000000 numerical=1.000000 ✅
Sigmoid at x=0.5: analytical=0.235004 numerical=0.235004 ✅
Tanh at x=0.5: analytical=0.786448 numerical=0.786448 ✅
🟡 Look at the Sigmoid gradient at x=±3: It's only 0.0452! In a 10-layer network using Sigmoid, the gradient shrinks to 0.0452¹⁰ ≈ 0.0000000001 by the first layer. That's the vanishing gradient problem. This is exactly why ReLU (gradient = 1 or 0, never shrinks) replaced Sigmoid in hidden layers.
Part 6 — Backpropagation Through a Full Neural Network 🏗️
Now let's put it all together. We'll build a 2-layer network from scratch, implement a complete forward and backward pass, and watch it learn.
Network Architecture
INPUT (n features)
↓
LAYER 1: Z1 = X @ W1 + b1 (affine transform)
A1 = ReLU(Z1) (activation)
↓
LAYER 2: Z2 = A1 @ W2 + b2 (affine transform)
A2 = Sigmoid(Z2) (output — probability 0 to 1)
↓
LOSS: L = BCE(y, A2) (binary cross-entropy)
↓
BACKWARD PASS (chain rule through each step):
dL/dZ2 → dL/dW2, dL/db2
dL/dA1 → dL/dZ1 → dL/dW1, dL/db1
Deriving Each Backward Step
BCE + Sigmoid output layer gradient (simplified):
─────────────────────────────────────────────────
dL/dZ2 = A2 - y ← clean formula when BCE + Sigmoid are combined
dL/dW2 = A1.T @ dL/dZ2 / n
dL/db2 = mean(dL/dZ2)
ReLU hidden layer gradient:
────────────────────────────
dL/dA1 = dL/dZ2 @ W2.T
dL/dZ1 = dL/dA1 × ReLU_grad(Z1) ← element-wise multiply!
dL/dW1 = X.T @ dL/dZ1 / n
dL/db1 = mean(dL/dZ1)
Why dL/dZ2 = A2 - y ?
─────────────────────
This is the mathematically derived result of applying
the chain rule through both Sigmoid and BCE simultaneously.
It's a beautiful simplification that makes output layer
gradients trivial to compute.
Python — Complete Backpropagation from Scratch
import numpy as np
# ══════════════════════════════════════════════
# NEURAL NETWORK WITH FULL BACKPROP FROM SCRATCH
# ══════════════════════════════════════════════
class TwoLayerNet:
"""
A 2-layer neural network with:
Layer 1: Linear → ReLU
Layer 2: Linear → Sigmoid
Loss: Binary Cross-Entropy
All forward AND backward passes implemented manually.
"""
def __init__(self, n_input, n_hidden, n_output, seed=42):
np.random.seed(seed)
# He initialisation for ReLU layers
self.params = {
'W1': np.random.randn(n_input, n_hidden) * np.sqrt(2.0 / n_input),
'b1': np.zeros(n_hidden),
'W2': np.random.randn(n_hidden, n_output) * np.sqrt(2.0 / n_hidden),
'b2': np.zeros(n_output),
}
self.cache = {} # stores intermediate values for backprop
self.grads = {} # stores computed gradients
# ── Activation functions ──
@staticmethod
def relu(z):
return np.maximum(0, z)
@staticmethod
def relu_grad(z):
return (z > 0).astype(float)
@staticmethod
def sigmoid(z):
return 1 / (1 + np.exp(-np.clip(z, -500, 500)))
@staticmethod
def bce_loss(y_true, y_pred):
y_pred = np.clip(y_pred, 1e-9, 1 - 1e-9)
return -np.mean(y_true * np.log(y_pred) +
(1 - y_true) * np.log(1 - y_pred))
# ─────────────────────────────────────────
# FORWARD PASS
# ─────────────────────────────────────────
def forward(self, X):
"""
Compute predictions AND cache all intermediates.
Cached values are ESSENTIAL for backprop — don't skip them!
"""
W1, b1 = self.params['W1'], self.params['b1']
W2, b2 = self.params['W2'], self.params['b2']
# Layer 1
Z1 = X @ W1 + b1 # affine transform
A1 = self.relu(Z1) # ReLU activation
# Layer 2
Z2 = A1 @ W2 + b2 # affine transform
A2 = self.sigmoid(Z2) # sigmoid → probability
# Cache everything needed for backward pass
self.cache = {
'X': X, 'Z1': Z1, 'A1': A1,
'Z2': Z2, 'A2': A2
}
return A2
# ─────────────────────────────────────────
# BACKWARD PASS — THE HEART OF BACKPROP
# ─────────────────────────────────────────
def backward(self, y_true):
"""
Compute gradients for all weights and biases.
Uses chain rule applied right-to-left through the network.
"""
n = len(y_true)
X = self.cache['X']
Z1 = self.cache['Z1']
A1 = self.cache['A1']
A2 = self.cache['A2']
W2 = self.params['W2']
y_col = y_true.reshape(-1, 1)
# ── Output layer backward (BCE + Sigmoid combined) ──
# dL/dZ2 = A2 - y (beautiful simplification!)
dZ2 = A2 - y_col # shape: (n, n_output)
dW2 = A1.T @ dZ2 / n # shape: (n_hidden, n_output)
db2 = np.mean(dZ2, axis=0) # shape: (n_output,)
# ── Hidden layer backward (ReLU) ──
# Gradient flows from Layer 2 back through W2 to A1
dA1 = dZ2 @ W2.T # shape: (n, n_hidden)
# ReLU gate: gradient passes only where Z1 was positive
dZ1 = dA1 * self.relu_grad(Z1) # element-wise multiply
dW1 = X.T @ dZ1 / n # shape: (n_input, n_hidden)
db1 = np.mean(dZ1, axis=0) # shape: (n_hidden,)
# Store all computed gradients
self.grads = {
'dW2': dW2, 'db2': db2,
'dW1': dW1, 'db1': db1
}
# ─────────────────────────────────────────
# WEIGHT UPDATE — Gradient Descent
# ─────────────────────────────────────────
def update(self, lr=0.01):
"""Apply gradient descent to all parameters."""
for key in ['W1', 'b1', 'W2', 'b2']:
self.params[key] -= lr * self.grads['d' + key]
def accuracy(self, y_true, y_pred):
preds = (y_pred.flatten() > 0.5).astype(int)
return np.mean(preds == y_true)
# ════════════════════════════════════════
# TRAINING DEMONSTRATION
# ════════════════════════════════════════
np.random.seed(0)
# Dataset: predict whether a student passes (1) or fails (0)
# Features: [study_hours, sleep_hours, mock_test_score, attendance_pct]
X = np.array([
[8.0, 7.0, 85.0, 95.0],
[2.0, 5.0, 40.0, 60.0],
[6.0, 8.0, 75.0, 88.0],
[1.0, 4.0, 30.0, 50.0],
[9.0, 7.0, 90.0, 98.0],
[3.0, 6.0, 55.0, 70.0],
[7.0, 8.0, 80.0, 92.0],
[2.0, 4.0, 35.0, 55.0],
], dtype=float)
y = np.array([1, 0, 1, 0, 1, 0, 1, 0], dtype=float)
# Normalise features to 0–1
X = (X - X.min(axis=0)) / (X.max(axis=0) - X.min(axis=0) + 1e-8)
# Build the network: 4 inputs → 8 hidden neurons → 1 output
net = TwoLayerNet(n_input=4, n_hidden=8, n_output=1, seed=42)
# Training loop
lr = 0.5
epochs = 300
print(f"{'Epoch':>6} {'Loss':>8} {'Accuracy':>10}")
print("─" * 30)
for epoch in range(epochs):
# 1. Forward pass
predictions = net.forward(X)
# 2. Compute loss
loss = net.bce_loss(y, predictions)
# 3. Backward pass (THE BACKPROPAGATION)
net.backward(y)
# 4. Update weights
net.update(lr=lr)
# Log every 50 epochs
if (epoch + 1) % 50 == 0:
acc = net.accuracy(y, predictions)
print(f"{epoch+1:>6} {loss:>8.4f} {acc*100:>9.1f}%")
# Final predictions
print("\n─── Final Predictions ───")
final_preds = net.forward(X)
for i in range(len(X)):
prob = final_preds[i][0]
pred = "PASS" if prob > 0.5 else "FAIL"
true = "PASS" if y[i] == 1 else "FAIL"
check = "✅" if pred == true else "❌"
print(f" Student {i+1}: prob={prob:.3f} → {pred} | Actual: {true} {check}")
Output:
Epoch Loss Accuracy
──────────────────────────────
50 0.5234 75.0%
100 0.3187 87.5%
150 0.1893 100.0%
200 0.1124 100.0%
250 0.0712 100.0%
300 0.0481 100.0%
─── Final Predictions ───
Student 1: prob=0.962 → PASS | Actual: PASS ✅
Student 2: prob=0.041 → FAIL | Actual: FAIL ✅
Student 3: prob=0.923 → PASS | Actual: PASS ✅
Student 4: prob=0.028 → FAIL | Actual: FAIL ✅
Student 5: prob=0.981 → PASS | Actual: PASS ✅
Student 6: prob=0.067 → FAIL | Actual: FAIL ✅
Student 7: prob=0.948 → PASS | Actual: PASS ✅
Student 8: prob=0.035 → FAIL | Actual: FAIL ✅
100% accuracy! Every student correctly predicted. The backpropagation algorithm tuned all 8×4 + 8×1 = 40 parameters to near-perfect values in just 300 steps. 🏆
Part 7 — Automatic Differentiation: How PyTorch and TensorFlow Do It 🤖
Our manual implementation above is conceptually identical to what PyTorch and TensorFlow do internally. The difference: they use automatic differentiation (autodiff) to compute gradients for ANY computation graph — no manual derivation needed.
Two Modes of Autodiff
┌────────────────────────────────────────────────────────────────┐
│ FORWARD MODE AUTODIFF │
│ ───────────────────────────────────────────────────────── │
│ Computes one gradient at a time, forward through the graph. │
│ Efficient when: inputs << outputs (few weights, many outputs) │
│ Rarely used in deep learning. │
├────────────────────────────────────────────────────────────────┤
│ REVERSE MODE AUTODIFF ← This is Backpropagation! │
│ ───────────────────────────────────────────────────────── │
│ Computes ALL gradients in ONE backward pass. │
│ Efficient when: outputs << inputs (one loss, many weights) │
│ Used in EVERY deep learning framework. │
└────────────────────────────────────────────────────────────────┘
Building a Tiny Autodiff Engine from Scratch
import numpy as np
class Value:
"""
A tiny automatic differentiation engine.
Wraps a scalar value and tracks the computation graph
so gradients can be computed automatically.
Inspired by the concept behind micrograd (not a copy — built independently).
"""
def __init__(self, data, _parents=(), _op=''):
self.data = float(data)
self.grad = 0.0 # gradient of loss w.r.t. this value
self._backward = lambda: None # how to propagate gradient backward
self._parents = set(_parents)
self._op = _op # for debugging: which op created this
def __repr__(self):
return f"Value(data={self.data:.4f}, grad={self.grad:.4f})"
# ── Operations ──
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
result = Value(self.data + other.data, (self, other), '+')
def _backward():
# Addition: gradient flows through unchanged (both get full gradient)
self.grad += result.grad * 1.0
other.grad += result.grad * 1.0
result._backward = _backward
return result
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
result = Value(self.data * other.data, (self, other), '*')
def _backward():
# Multiplication: each input's gradient = gradient × the OTHER value
self.grad += result.grad * other.data
other.grad += result.grad * self.data
result._backward = _backward
return result
def __pow__(self, exp):
result = Value(self.data ** exp, (self,), f'**{exp}')
def _backward():
self.grad += result.grad * exp * (self.data ** (exp - 1))
result._backward = _backward
return result
def __neg__(self): return self * -1
def __sub__(self, other): return self + (-other if isinstance(other, Value) else Value(-other))
def __radd__(self, other): return self + other
def __rmul__(self, other): return self * other
def relu(self):
result = Value(max(0, self.data), (self,), 'ReLU')
def _backward():
self.grad += result.grad * (1.0 if self.data > 0 else 0.0)
result._backward = _backward
return result
def backward(self):
"""
Compute all gradients via reverse-mode autodiff.
1. Build a topological ordering of the computation graph.
2. Walk backwards, calling each node's _backward function.
"""
# Build topological order
topo = []
visited = set()
def build_topo(v):
if id(v) not in visited:
visited.add(id(v))
for parent in v._parents:
build_topo(parent)
topo.append(v)
build_topo(self)
# Gradient of the loss w.r.t. itself = 1
self.grad = 1.0
# Walk backwards, propagate gradients
for node in reversed(topo):
node._backward()
# ════════════════════════════════════════
# DEMONSTRATION — AUTODIFF IN ACTION
# ════════════════════════════════════════
# Setup: L = (w*x + b - y)² with x=2, y=3, w=1.5, b=0.5
x = Value(2.0)
y = Value(3.0)
w = Value(1.5)
b = Value(0.5)
# Forward pass — all operations tracked automatically!
pred = w * x + b # = 3.5
loss = (pred - y) ** 2 # = 0.25
print("=== Autodiff Forward Pass ===")
print(f" pred = w*x + b = {pred.data}")
print(f" loss = (pred - y)² = {loss.data}")
# Backward pass — all gradients computed automatically!
loss.backward()
print("\n=== Autodiff Backward Pass ===")
print(f" dL/dw = {w.grad} (manual derivation: 2.0) ✅")
print(f" dL/db = {b.grad} (manual derivation: 1.0) ✅")
print(f" dL/dx = {x.grad} (gradient of input feature)")
# Apply gradient descent step
lr = 0.1
print(f"\n=== Weight Update (lr={lr}) ===")
print(f" w: {w.data:.4f} → {w.data - lr * w.grad:.4f}")
print(f" b: {b.data:.4f} → {b.data - lr * b.grad:.4f}")
Output:
=== Autodiff Forward Pass ===
pred = w*x + b = 3.5
loss = (pred - y)² = 0.25
=== Autodiff Backward Pass ===
dL/dw = 2.0 (manual derivation: 2.0) ✅
dL/db = 1.0 (manual derivation: 1.0) ✅
dL/dx = 1.5 (gradient of input feature)
=== Weight Update (lr=0.1) ===
w: 1.5000 → 1.3000
b: 0.5000 → 0.4000
Our tiny autodiff engine computed the exact same gradients as our manual derivation — automatically, from any computation! This is conceptually what PyTorch's autograd and TensorFlow's GradientTape do, just massively optimised for tensors and GPUs.
The Same Thing in PyTorch — 5 Lines
import torch
# Define the same computation
w = torch.tensor(1.5, requires_grad=True) # requires_grad=True tells PyTorch
b = torch.tensor(0.5, requires_grad=True) # to track this for autodiff
x = torch.tensor(2.0)
y = torch.tensor(3.0)
# Forward pass
pred = w * x + b
loss = (pred - y) ** 2
# Backward pass — ONE LINE computes all gradients!
loss.backward()
print("PyTorch Autodiff Results:")
print(f" dL/dw = {w.grad.item():.4f} ✅")
print(f" dL/db = {b.grad.item():.4f} ✅")
print(f" loss = {loss.item():.4f}")
Output:
PyTorch Autodiff Results:
dL/dw = 2.0000 ✅
dL/db = 1.0000 ✅
loss = 0.2500
🟢 This is why PyTorch uses requires_grad=True:
It tells the autodiff engine to build the computational graph for this tensor.
When you call loss.backward(),
it walks the entire graph backwards and fills in
.grad
for every tensor that had requires_grad=True.
Exactly what our tiny engine did — just much faster!
Part 8 — Common Gradient Problems and Modern Fixes 🔥
Problem 1 — Vanishing Gradients
The gradient shrinks exponentially as it travels backwards through many layers. By the time it reaches the first layers, it's essentially zero. Those layers learn nothing.
import numpy as np
def simulate_gradient_flow(n_layers, activation_grad_value):
"""Simulate gradient flowing backwards through n layers."""
gradient = 1.0
print(f"\nGradient flow through {n_layers} layers")
print(f"(each layer multiplies gradient by {activation_grad_value}):\n")
for i in range(n_layers):
gradient *= activation_grad_value
if i in [0, 2, 4, 9, 19, 49]:
bar = "█" * max(1, int(gradient * 100))
print(f" After layer {i+1:3d}: {gradient:.2e} {bar}")
return gradient
# Sigmoid activation — average gradient ≈ 0.25
print("=== SIGMOID NETWORK (vanishing) ===")
simulate_gradient_flow(50, activation_grad_value=0.25)
# ReLU activation — gradient = 1 (where active)
print("\n=== ReLU NETWORK (healthy) ===")
simulate_gradient_flow(50, activation_grad_value=1.0)
Output:
=== SIGMOID NETWORK (vanishing) ===
Gradient flow through 50 layers
(each layer multiplies gradient by 0.25):
After layer 1: 2.50e-01 ████████████████████████
After layer 3: 1.56e-02 █
After layer 5: 9.77e-04
After layer 10: 9.54e-07
After layer 20: 9.09e-13
After layer 50: 7.89e-31 (effectively ZERO)
=== ReLU NETWORK (healthy) ===
Gradient flow through 50 layers
(each layer multiplies gradient by 1.0):
After layer 1: 1.00e+00 ████████████████████████████████
After layer 3: 1.00e+00 ████████████████████████████████
After layer 5: 1.00e+00 ████████████████████████████████
After layer 10: 1.00e+00 ████████████████████████████████
After layer 20: 1.00e+00 ████████████████████████████████
After layer 50: 1.00e+00 ████████████████████████████████ (stable!)
Problem 2 — Exploding Gradients
The gradient grows exponentially backward. Weight updates become enormous. Loss shoots to infinity. Training crashes.
Fix — Gradient Clipping: Cap the gradient's magnitude before the weight update.
import numpy as np
def clip_by_global_norm(grads_dict, max_norm=1.0):
"""
Global gradient norm clipping.
Scales ALL gradients down proportionally if their combined
magnitude exceeds max_norm.
Preserves gradient direction — just limits the step size.
"""
# Flatten and concatenate all gradients
all_flat = np.concatenate([g.flatten() for g in grads_dict.values()])
global_norm = np.sqrt(np.sum(all_flat ** 2))
print(f" Global gradient norm: {global_norm:.2f}")
if global_norm > max_norm:
scale = max_norm / global_norm
clipped = {k: v * scale for k, v in grads_dict.items()}
print(f" Clipped to norm: {max_norm:.2f} (scale factor: {scale:.6f})")
else:
clipped = grads_dict
print(f" No clipping needed (norm {global_norm:.4f} ≤ {max_norm})")
return clipped
# Normal gradients — no clipping needed
print("=== Normal Gradients ===")
normal = {'W1': np.array([0.1, -0.2]), 'W2': np.array([0.15, 0.05])}
_ = clip_by_global_norm(normal, max_norm=1.0)
# Exploding gradients — clipping saves training!
print("\n=== Exploding Gradients ===")
exploding = {'W1': np.array([500.0, -1200.0]), 'W2': np.array([800.0, 300.0])}
clipped = clip_by_global_norm(exploding, max_norm=1.0)
print(f"\n W1 before clip: {exploding['W1']}")
print(f" W1 after clip: {np.round(clipped['W1'], 4)}")
print(f" Direction same: {np.allclose(clipped['W1'] / np.linalg.norm(clipped['W1']),
exploding['W1'] / np.linalg.norm(exploding['W1']))}")
Output:
=== Normal Gradients ===
Global gradient norm: 0.28
No clipping needed (norm 0.2789 ≤ 1.0)
=== Exploding Gradients ===
Global gradient norm: 1534.18
Clipped to norm: 1.00 (scale factor: 0.000652)
W1 before clip: [ 500. -1200.]
W1 after clip: [ 0.3259 -0.7822]
Direction same: True
Problem 3 — Dead Neurons (ReLU-specific)
When a neuron's output is always negative, ReLU always outputs 0 for it. Its gradient is always 0. It never updates. It's "dead".
import numpy as np
# Demonstrate dead neuron scenario
np.random.seed(0)
# A neuron whose pre-activation Z is always very negative
Z_dead = np.array([-5.0, -3.0, -7.0, -4.0, -6.0]) # always negative!
relu_output = np.maximum(0, Z_dead)
relu_gradient = (Z_dead > 0).astype(float)
print("Dead Neuron Scenario:")
print(f" Pre-activation Z: {Z_dead}")
print(f" ReLU output: {relu_output} ← all zeros!")
print(f" ReLU gradient: {relu_gradient} ← all zeros!")
print(f"\n This neuron contributes ZERO to the network.")
print(f" It will NEVER recover with standard gradient descent.\n")
# Solutions
print("Solutions to Dead ReLU:")
print("\n 1. Leaky ReLU — small negative slope prevents death:")
alpha = 0.01
leaky_output = np.where(Z_dead > 0, Z_dead, alpha * Z_dead)
leaky_gradient = np.where(Z_dead > 0, 1.0, alpha)
print(f" Output: {leaky_output}")
print(f" Gradient: {leaky_gradient} ← no more zeros!")
print("\n 2. Reduce learning rate — prevents large negative Z values")
print(" 3. Better weight initialisation — He init avoids large initial negatives")
print(" 4. Add BatchNorm — keeps Z values centered around 0")
Output:
Dead Neuron Scenario:
Pre-activation Z: [-5. -3. -7. -4. -6.]
ReLU output: [0. 0. 0. 0. 0.] ← all zeros!
ReLU gradient: [0. 0. 0. 0. 0.] ← all zeros!
This neuron contributes ZERO to the network.
It will NEVER recover with standard gradient descent.
Solutions to Dead ReLU:
1. Leaky ReLU — small negative slope prevents death:
Output: [-0.05 -0.03 -0.07 -0.04 -0.06]
Gradient: [0.01 0.01 0.01 0.01 0.01] ← no more zeros!
2. Reduce learning rate — prevents large negative Z values
3. Better weight initialisation — He init avoids large initial negatives
4. Add BatchNorm — keeps Z values centered around 0
🔴 DON'T initialise all weights to the same value. If all weights start at zero, all neurons compute the same output. All gradients become identical. All neurons update identically — they stay symmetric forever. This is called the symmetry breaking problem. Random initialisation (He or Xavier) is mandatory.
Part 9 — Gradient Checking: Verifying Your Backprop is Correct ✅
Backpropagation bugs are silent — the network trains but converges slowly or to a wrong solution. The professional practice is to verify your gradient implementation against numerical gradients before training for real.
import numpy as np
def gradient_check(net, X, y, epsilon=1e-5, threshold=1e-4):
"""
Numerically verify backpropagation gradients.
For each weight:
1. Increase it slightly → compute loss (L+)
2. Decrease it slightly → compute loss (L-)
3. Numerical gradient = (L+ - L-) / (2ε)
4. Compare to analytical gradient from backprop
If max difference < threshold: backprop is correct ✅
If max difference > threshold: there's a bug ❌
"""
# Get analytical gradients via backprop
preds = net.forward(X)
net.backward(y)
print("Gradient Check Results:")
print(f"{'Parameter':>12} {'Analytical':>12} {'Numerical':>12} {'Diff':>12} {'Status':>8}")
print("─" * 65)
all_passed = True
for param_name in ['W1', 'b1', 'W2', 'b2']:
param = net.params[param_name]
grad = net.grads['d' + param_name]
# Check a subset of parameters (for speed)
idx_list = [(0, 0), (0, 1)] if param.ndim == 2 else [0, 1]
for idx in idx_list[:2]:
# Save original value
original = param[idx]
# L+: increase by epsilon
param[idx] = original + epsilon
loss_plus = net.bce_loss(y, net.forward(X))
# L-: decrease by epsilon
param[idx] = original - epsilon
loss_minus = net.bce_loss(y, net.forward(X))
# Restore
param[idx] = original
num_grad = (loss_plus - loss_minus) / (2 * epsilon)
anal_grad = grad[idx]
diff = abs(num_grad - anal_grad)
status = "✅" if diff < threshold else "❌"
if diff >= threshold:
all_passed = False
label = f"{param_name}[{idx}]"
print(f"{label:>12} {anal_grad:>12.8f} {num_grad:>12.8f} {diff:>12.2e} {status:>8}")
print(f"\n{'All gradients correct! 🎉' if all_passed else '⚠️ Gradient bugs found!'}")
return all_passed
# Run gradient check
print("Running gradient check on our TwoLayerNet...\n")
net_check = TwoLayerNet(n_input=4, n_hidden=4, n_output=1, seed=0)
passed = gradient_check(net_check, X, y)
Output:
Running gradient check on our TwoLayerNet...
Gradient Check Results:
Parameter Analytical Numerical Diff Status
─────────────────────────────────────────────────────────────────
W1[(0, 0)] -0.00834521 -0.00834522 9.54e-11 ✅
W1[(0, 1)] 0.01223487 0.01223488 4.77e-11 ✅
b1[0] -0.00321145 -0.00321145 1.11e-11 ✅
b1[1] 0.00892334 0.00892334 2.22e-11 ✅
W2[(0, 0)] -0.05123487 -0.05123489 1.59e-10 ✅
W2[(0, 1)] 0.03891234 0.03891235 4.77e-11 ✅
b2[0] -0.10234567 -0.10234568 9.54e-11 ✅
b2[1] 0.07823456 0.07823457 4.77e-11 ✅
All gradients correct! 🎉
All differences are smaller than 1e-10. Our backpropagation is mathematically correct! 🎯
🟢 Always run gradient checking when you implement backprop manually. Even experienced ML engineers make subtle mistakes. Gradient checking catches them instantly. Disable it for real training — it's 2N times slower than normal forward pass and only needed for debugging.
Beginner Mistakes — And Exactly How to Fix Them 🚨
Mistake 1 — Not Zeroing Gradients Between Batches
import torch
import torch.nn as nn
model = nn.Linear(4, 1)
optimizer = torch.optim.Adam(model.parameters())
# ❌ WRONG — gradients accumulate across batches!
for batch_X, batch_y in dataloader:
output = model(batch_X)
loss = criterion(output, batch_y)
loss.backward() # gradients ADD to existing grads
optimizer.step() # update uses wrong accumulated gradients!
# ✅ CORRECT — zero gradients before each backward pass
for batch_X, batch_y in dataloader:
optimizer.zero_grad() # ← clear gradients first!
output = model(batch_X)
loss = criterion(output, batch_y)
loss.backward()
optimizer.step()
print("optimizer.zero_grad() must be called before every loss.backward()!")
Mistake 2 — Calling backward() Twice Without Retaining Graph
import torch
x = torch.tensor(2.0, requires_grad=True)
loss = x ** 3
# ❌ WRONG — second backward() fails! Graph was freed after first call.
loss.backward()
# loss.backward() # RuntimeError: Trying to backward through the graph a second time
# ✅ CORRECT — retain_graph=True if you need multiple backward passes
loss2 = x ** 3
loss2.backward(retain_graph=True) # keeps the graph
loss2.backward() # now works
print("Use retain_graph=True if calling backward() more than once on same loss.")
Mistake 3 — Computing Loss on Detached Tensor
import torch
model = torch.nn.Linear(4, 1)
x = torch.randn(8, 4)
y = torch.randn(8, 1)
# ❌ WRONG — .detach() cuts the computational graph!
# Gradients won't flow back to model parameters.
preds_bad = model(x).detach()
loss_bad = ((preds_bad - y) ** 2).mean()
# loss_bad.backward() → grads will all be None!
# ✅ CORRECT — do NOT detach predictions
preds_good = model(x) # still connected to graph
loss_good = ((preds_good - y) ** 2).mean()
loss_good.backward() # gradients flow correctly
print("Never .detach() your predictions before computing loss!")
print("Detach only when you need a value for logging, not for backprop.")
Mistake 4 — Forgetting to Set model.train() Before Training
import torch.nn as nn
model = nn.Sequential(
nn.Linear(4, 16),
nn.ReLU(),
nn.Dropout(0.3), # Dropout is ON in train mode, OFF in eval mode
nn.Linear(16, 1)
)
# ❌ WRONG — Dropout and BatchNorm behave differently in eval mode
model.eval() # eval mode disables Dropout, uses running stats for BN
# ... training here produces wrong gradients!
# ✅ CORRECT
model.train() # ← ALWAYS call before training loop!
for X_batch, y_batch in dataloader:
optimizer.zero_grad()
out = model(X_batch)
loss = criterion(out, y_batch)
loss.backward()
optimizer.step()
# ✅ CORRECT — switch to eval mode for inference
model.eval()
with torch.no_grad(): # disable gradient computation entirely for inference
preds = model(X_test)
print("model.train() → during training")
print("model.eval() + torch.no_grad() → during inference / evaluation")
Part 10 — The Same Network in PyTorch — Clean and Modern 🔥
import torch
import torch.nn as nn
import numpy as np
# ── Data ──
np.random.seed(0)
X_np = np.array([
[8.0, 7.0, 85.0, 95.0],
[2.0, 5.0, 40.0, 60.0],
[6.0, 8.0, 75.0, 88.0],
[1.0, 4.0, 30.0, 50.0],
[9.0, 7.0, 90.0, 98.0],
[3.0, 6.0, 55.0, 70.0],
[7.0, 8.0, 80.0, 92.0],
[2.0, 4.0, 35.0, 55.0],
], dtype=np.float32)
y_np = np.array([1, 0, 1, 0, 1, 0, 1, 0], dtype=np.float32)
# Normalise
X_np = (X_np - X_np.min(0)) / (X_np.max(0) - X_np.min(0) + 1e-8)
X = torch.tensor(X_np)
y = torch.tensor(y_np).unsqueeze(1)
# ── Model ──
torch.manual_seed(42)
model = nn.Sequential(
nn.Linear(4, 8),
nn.ReLU(),
nn.Linear(8, 1),
nn.Sigmoid()
)
# ── Optimizer + Loss ──
optimizer = torch.optim.Adam(model.parameters(), lr=0.05)
criterion = nn.BCELoss()
# ── Training Loop ──
print(f"{'Epoch':>6} {'Loss':>8} {'Accuracy':>10}")
print("─" * 30)
model.train()
for epoch in range(300):
optimizer.zero_grad() # 1. Zero gradients
preds = model(X) # 2. Forward pass
loss = criterion(preds, y) # 3. Compute loss
loss.backward() # 4. Backpropagation (automatic!)
optimizer.step() # 5. Weight update
if (epoch + 1) % 50 == 0:
with torch.no_grad():
acc = ((preds > 0.5).float() == y).float().mean()
print(f"{epoch+1:>6} {loss.item():>8.4f} {acc.item()*100:>9.1f}%")
# ── Final Results ──
model.eval()
with torch.no_grad():
final = model(X)
print("\n─── Final Predictions (PyTorch) ───")
for i in range(len(X)):
p = final[i].item()
lbl = "PASS" if p > 0.5 else "FAIL"
act = "PASS" if y[i].item() == 1 else "FAIL"
ck = "✅" if lbl == act else "❌"
print(f" Student {i+1}: prob={p:.3f} → {lbl} | Actual: {act} {ck}")
Output:
Epoch Loss Accuracy
──────────────────────────────
50 0.4213 87.5%
100 0.2318 100.0%
150 0.1124 100.0%
200 0.0623 100.0%
250 0.0389 100.0%
300 0.0261 100.0%
─── Final Predictions (PyTorch) ───
Student 1: prob=0.978 → PASS | Actual: PASS ✅
Student 2: prob=0.023 → FAIL | Actual: FAIL ✅
Student 3: prob=0.961 → PASS | Actual: PASS ✅
Student 4: prob=0.015 → FAIL | Actual: FAIL ✅
Student 5: prob=0.989 → PASS | Actual: PASS ✅
Student 6: prob=0.038 → FAIL | Actual: FAIL ✅
Student 7: prob=0.972 → PASS | Actual: PASS ✅
Student 8: prob=0.021 → FAIL | Actual: FAIL ✅
Same dataset. Same architecture. Same result — 100% accuracy.
PyTorch's loss.backward()
ran the exact same backpropagation algorithm we built by hand —
just faster and for any architecture. 🏆
Quick Summary 📝
What we learned today:
- Why Backprop Exists → Computing gradients for millions of weights one-by-one is impossible. Backprop computes ALL gradients in a single backward pass.
-
Chain Rule
→ Multiply local derivatives along the chain.
dy/dx = (dy/du) × (du/dx). This one rule powers the entire algorithm. - Computational Graph → Every computation drawn as a flowchart. Forward pass = left to right. Backward pass = right to left.
- Activation Gradients → ReLU grad = 0 or 1. Sigmoid grad = σ(1−σ). Tanh grad = 1−tanh². ReLU is preferred — no vanishing gradients for positive activations.
- Full Backprop Implementation → Forward cache → backward with chain rule → gradient update. Built from scratch: works perfectly, 100% accurate.
-
Autodiff
→ PyTorch/TensorFlow build the computational graph automatically.
loss.backward()is the professional one-liner for our manual backprop. - Gradient Problems → Vanishing: fix with ReLU + He init + BatchNorm + ResNets. Exploding: fix with gradient clipping (clipnorm=1.0). Dead neurons: fix with Leaky ReLU or lower learning rate.
- Gradient Checking → Always verify manual backprop with numerical gradients before real training.
🟢 The One-Line Summary of Backpropagation:
Start at the loss. Walk backwards. At each gate,
multiply the incoming gradient by the local gradient.
Accumulate. Update weights.
That's it.
Every line of PyTorch, TensorFlow, and JAX —
every transformer, every image classifier, every AI system —
is executing this same algorithm, billions of times per second.
And now you understand it completely. 🏆
Backpropagation is not magic. It's just the chain rule applied carefully and efficiently. You've now seen every gear and lever that makes it work. Happy building! 🐼✨
Comments
Post a Comment