The deep learning training loop is the repeatable process through which a neural network predicts, measures error, calculates a correction, updates its parameters, and learns from the next batch of evidence. 🏀
It matters because a model can appear to learn while actually memorising noisy labels, leaking future information, or becoming too slow and costly to deploy. A reliable training loop combines mathematics with disciplined data, validation, monitoring, and release controls. 📈
📑 In This Post
🔀 Quick Comparison
| Stage | Primary question | Key evidence |
|---|---|---|
| Training | Can parameters reduce the defined loss? | Training loss and numerical stability. |
| Validation | Does the model generalise to held-out data? | Task metrics, slice analysis, calibration. |
| Production | Does it remain useful under live conditions? | Drift, outcomes, latency, throughput, cost. |
The Training Loop — Big Picture 🗺️
Here is the complete journey a neural network takes every single training step:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
┌─────────────────────────────────────────────────────────────┐
│ THE TRAINING LOOP │
│ │
│ 📥 Input Data │
│ ↓ │
│ 🔮 Forward Pass (network makes a prediction) │
│ ↓ │
│ 📊 Predictions (what the network guessed) │
│ ↓ │
│ 📏 Loss Calculation (how wrong was the guess?) │
│ ↓ │
│ 🔄 Backpropagation (trace WHERE the mistake came from)│
│ ↓ │
│ 🔧 Weight Update (fix the mistake a tiny bit) │
│ ↓ │
│ 🔁 Repeat (do it all again, thousands of │
│ times, until good enough) │
└─────────────────────────────────────────────────────────────┘
Each step in this loop has a specific job. Together they teach the network to go from "I have no idea what this is" to "That is definitely a cat."
Step 1 — Input Data: Feeding the Network 📥
Before a student can answer a question, the teacher has to show them the question.
Input data is exactly that — the question shown to the network. It could be pixels from an image, numbers from a sensor, words from a sentence, or any other information.
What Does Input Data Look Like in Code?
Every piece of input is a matrix of numbers. A batch of inputs is called X (by convention). The correct answers for those inputs are called y.
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
import numpy as np
# ── Toy Dataset: Predict if a student passes (1) or fails (0) ──
# Features: [hours_studied, hours_slept, practice_tests_done]
X = np.array([
[2.0, 6.0, 1.0], # Student 1 → barely studied
[5.0, 7.0, 3.0], # Student 2 → moderate prep
[8.0, 8.0, 5.0], # Student 3 → well prepared
[1.0, 4.0, 0.0], # Student 4 → very under-prepared
[7.0, 7.0, 4.0], # Student 5 → well prepared
[3.0, 5.0, 2.0], # Student 6 → average
])
# Labels: 1 = pass, 0 = fail
y = np.array([0, 1, 1, 0, 1, 0])
print("Input shape (X):", X.shape) # 6 students, 3 features each
print("Labels shape (y):", y.shape) # 6 answers
print("\nFirst student's features:", X[0])
print("First student's label:", y[0], "(fail)")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Input shape (X): (6, 3)
Labels shape (y): (6,)
First student's features: [2. 6. 1.]
First student's label: 0 (fail)
What is a Batch?
Training on the whole dataset at once would be very slow. Instead, we split data into small chunks called batches.
- Batch size = 32 → process 32 samples, update weights, then next 32
- One pass through all batches = one epoch
- Training for 50 epochs = 50 full sweeps through the entire dataset
🟡 Batch size tip: Small batches (8–32) introduce randomness that actually helps the network escape bad solutions — but training is noisier. Large batches (256–1024) are faster and more stable but can sometimes converge to less optimal solutions. 32 or 64 is the sweet spot for most beginners.
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
def get_batches(X, y, batch_size):
"""
Split dataset into mini-batches.
Yields (X_batch, y_batch) tuples.
"""
n = len(X)
indices = np.random.permutation(n) # shuffle every epoch
X, y = X[indices], y[indices]
for start in range(0, n, batch_size):
end = start + batch_size
yield X[start:end], y[start:end]
# Preview what batching looks like
batch_size = 3
print(f"Dataset size: {len(X)} samples")
print(f"Batch size: {batch_size}")
print(f"Number of batches per epoch: {len(X) // batch_size}\n")
for i, (X_batch, y_batch) in enumerate(get_batches(X, y, batch_size)):
print(f"Batch {i+1}: X shape={X_batch.shape}, y shape={y_batch.shape}")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Dataset size: 6 samples
Batch size: 3
Number of batches per epoch: 2
Batch 1: X shape=(3, 3), y shape=(3,)
Batch 2: X shape=(3, 3), y shape=(3,)
Step 2 — Forward Pass: The Network Makes Its Guess 🔮
You show the network a photo of a cat. The data flows forward through the network — layer by layer — like water flowing down a stream.
Each layer transforms the data a little bit, until the last layer spits out a final answer: "I think there's a 92% chance this is a cat."
This one-way data journey — input to output — is the forward pass.
How Data Flows Through Layers
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
INPUT HIDDEN LAYER 1 HIDDEN LAYER 2 OUTPUT
[hours, → relu(W1 @ x + b1) → relu(W2 @ z1 + b2) → sigmoid(W3 @ z2 + b3)
sleep,
tests]
3 features 4 neurons 4 neurons 1 neuron (0 to 1)
Each → means: apply weights, add bias, apply activation
Python Code — Building and Running a Forward Pass
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
# ── Activation functions ──
def relu(z):
return np.maximum(0, z)
def sigmoid(z):
"""Squashes any number into range (0, 1) — perfect for probabilities."""
return 1 / (1 + np.exp(-z))
# ── Initialize weights for a 3-layer network ──
# Architecture: 3 inputs → 4 hidden → 4 hidden → 1 output
np.random.seed(42)
def he_init(n_in, n_out):
"""He initialization: best for ReLU layers."""
return np.random.randn(n_in, n_out) * np.sqrt(2.0 / n_in)
params = {
'W1': he_init(3, 4), 'b1': np.zeros(4),
'W2': he_init(4, 4), 'b2': np.zeros(4),
'W3': he_init(4, 1), 'b3': np.zeros(1),
}
# ── Forward pass function ──
def forward_pass(X, params):
"""
Pass input X through all layers.
Returns predictions AND all intermediate values
(needed for backpropagation later).
"""
cache = {}
# Layer 1: 3 inputs → 4 neurons + ReLU
cache['Z1'] = X @ params['W1'] + params['b1']
cache['A1'] = relu(cache['Z1'])
# Layer 2: 4 → 4 neurons + ReLU
cache['Z2'] = cache['A1'] @ params['W2'] + params['b2']
cache['A2'] = relu(cache['Z2'])
# Output layer: 4 → 1 + Sigmoid (gives probability 0–1)
cache['Z3'] = cache['A2'] @ params['W3'] + params['b3']
cache['A3'] = sigmoid(cache['Z3']) # final prediction
return cache['A3'], cache
# ── Run a forward pass on our student dataset ──
predictions, cache = forward_pass(X, params)
print("Input X shape:", X.shape)
print("Predictions shape:", predictions.shape)
print("\nRaw predictions (probabilities of passing):")
for i, pred in enumerate(predictions):
label = "PASS" if pred[0] > 0.5 else "FAIL"
print(f" Student {i+1}: {pred[0]:.4f} → Predicted: {label} | Actual: {'PASS' if y[i]==1 else 'FAIL'}")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Input X shape: (6, 3)
Predictions shape: (6, 1)
Raw predictions (probabilities of passing):
Student 1: 0.5021 → Predicted: PASS | Actual: FAIL
Student 2: 0.5019 → Predicted: PASS | Actual: PASS
Student 3: 0.5032 → Predicted: PASS | Actual: PASS
Student 4: 0.5013 → Predicted: PASS | Actual: FAIL
Student 5: 0.5028 → Predicted: PASS | Actual: PASS
Student 6: 0.5017 → Predicted: PASS | Actual: FAIL
The untrained network predicts roughly 0.50 for everyone — it's basically guessing! 🎲 This is completely normal. The training loop is about to fix this.
🟢 Key insight: A freshly initialised network has random weights — so its first predictions are essentially random. The entire purpose of the training loop is to push those random weights toward the right values so predictions become accurate.
Step 3 — Loss Calculation: Measuring How Wrong We Are 📏
After your basketball throw, your coach measures how far the ball landed from the hoop. 1 metre away? Pretty bad. 3 centimetres away? Almost perfect!
The loss (also called cost or error) is that measurement of wrongness for the neural network. A high loss = the network is very wrong. A low loss = the network is getting it right.
The goal of training is simple: make the loss as small as possible.
Common Loss Functions and When to Use Each
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
┌──────────────────────────────────────────────────────────────┐
│ CHOOSING YOUR LOSS FUNCTION │
│ │
│ PROBLEM TYPE LOSS FUNCTION OUTPUT │
│ ─────────────────────────────────────────────────────────── │
│ Binary Classification Binary Cross-Entropy sigmoid → 0/1 │
│ (pass/fail, spam/ham) │
│ │
│ Multi-Class Categorical softmax → one │
│ (cat/dog/bird) Cross-Entropy of N classes │
│ │
│ Regression Mean Squared Error any number │
│ (predict price/temp) (MSE) │
└──────────────────────────────────────────────────────────────┘
Binary Cross-Entropy Loss — Explained Simply
Our problem is binary (pass=1 or fail=0), so we use Binary Cross-Entropy (BCE).
The idea is elegant:
- If the true label is 1 (pass) and we predicted 0.95 → very small loss (good!)
- If the true label is 1 (pass) and we predicted 0.05 → very large loss (wrong!)
- If the true label is 0 (fail) and we predicted 0.05 → very small loss (good!)
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
BCE Formula (for one sample):
Loss = − [ y × log(ŷ) + (1−y) × log(1−ŷ) ]
Where:
y = true label (0 or 1)
ŷ = predicted probability (between 0 and 1)
For a batch: average the loss over all samples.
Python Code — Loss Calculation
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
def binary_cross_entropy(y_true, y_pred):
"""
Binary Cross-Entropy loss for binary classification.
y_true: shape (n,) or (n,1) — actual labels (0 or 1)
y_pred: shape (n,) or (n,1) — predicted probabilities (0.0 to 1.0)
Returns: single scalar — the average loss over all samples
"""
# Clip predictions to avoid log(0) which is undefined (−infinity)
y_pred = np.clip(y_pred, 1e-9, 1 - 1e-9)
# BCE formula
loss = -np.mean(
y_true * np.log(y_pred) +
(1 - y_true) * np.log(1 - y_pred)
)
return loss
def mean_squared_error(y_true, y_pred):
"""MSE loss for regression problems."""
return np.mean((y_true - y_pred) ** 2)
# ── Measure how wrong our untrained network is ──
y_pred = predictions.flatten() # from forward pass above
y_true = y.astype(float)
loss = binary_cross_entropy(y_true, y_pred)
print("True labels: ", y_true)
print("Predictions: ", np.round(y_pred, 4))
print(f"\nBinary Cross-Entropy Loss: {loss:.6f}")
print("\nInterpretation:")
print(" ~0.693 = completely random guessing (log(0.5))")
print(" ~0.000 = perfect predictions")
print(f" Our loss {loss:.3f} ≈ random ← expected for untrained network!")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
True labels: [0. 1. 1. 0. 1. 0.]
Predictions: [0.5021 0.5019 0.5032 0.5013 0.5028 0.5017]
Binary Cross-Entropy Loss: 0.693095
Interpretation:
~0.693 = completely random guessing (log(0.5))
~0.000 = perfect predictions
Our loss 0.693 ≈ random ← expected for untrained network!
Our loss of 0.693 confirms the network is guessing randomly. After training, we want to see this drop toward 0. 📉
Watching Loss Over Time — What It Should Look Like
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Loss
│
0.7│●
│ ●
0.5│ ●
│ ●
0.3│ ●
│ ● ●
0.1│ ● ● ● ● ●
│ ●●●●●
0.0│──────────────────────────────────── Epoch
0 10 20 30 40 50 60 70
✅ GOOD: Loss drops smoothly and levels off (converging)
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Loss
│
0.9│● ● ● ● ● ● ●
│ ● ● ● ● ● ● ●
0.5│
│
0.1│──────────────────────────────────── Epoch
❌ BAD: Loss bouncing wildly (learning rate too high)
🔴 DON'T ignore your loss curve. If the loss is not going down after 10–20 epochs, something is wrong — bad learning rate, wrong loss function, or data not normalised. Always plot your loss during training!
Step 4 — Backpropagation: Tracing Where the Mistake Came From 🔄
Your basketball missed by 2 metres to the left. Your coach asks: "Where did the mistake start?"
Was it your wrist? Your elbow? Your foot position? The coach traces backwards from the miss — release point → elbow angle → shoulder rotation → foot stance — and tells each body part: "You contributed this much to the miss."
Backpropagation does exactly the same thing. It starts from the loss and walks backwards through every layer, calculating how much each weight contributed to the mistake.
The Key Tool — The Chain Rule
Backpropagation uses one idea from calculus called the chain rule. You don't need to know calculus to understand the intuition:
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
Imagine a factory assembly line:
Raw material → Machine A → Machine B → Machine C → Final Product
If the final product is defective, you ask:
"How much did Machine C cause this?"
"How much did Machine B cause this?"
"How much did Machine A cause this?"
You go BACKWARDS through the line, calculating blame at each step.
That's the chain rule — chaining the blame backwards.
Visual Flow — Backprop Direction
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
FORWARD PASS (left to right):
──────────────────────────────────────────────────────────►
Input → Layer1 → Layer2 → Output → Loss
BACKWARD PASS (right to left):
◄──────────────────────────────────────────────────────────
Loss → ∂L/∂W3 → ∂L/∂W2 → ∂L/∂W1
∂L/∂W means: "How much does the loss change
when we change weight W by a tiny amount?"
This is called the GRADIENT of W.
Python Code — Backpropagation from Scratch
Let's implement backpropagation manually so you see exactly what's happening. No magic boxes — just math made visible.
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
def relu(z):
return np.maximum(0, z)
def sigmoid(z):
return 1 / (1 + np.exp(-z))
def relu_derivative(z):
"""
Gradient of ReLU:
1 where z > 0 (slope = 1, gradient passes through)
0 where z ≤ 0 (slope = 0, gradient is blocked)
"""
return (z > 0).astype(float)
def sigmoid_derivative(a):
"""
Gradient of sigmoid.
a = sigmoid(z) already computed.
"""
return a * (1 - a)
def backward_pass(X, y_true, predictions, cache, params):
"""
Compute gradients for all weights and biases
using backpropagation.
Returns dict of gradients: dW1, db1, dW2, db2, dW3, db3
"""
n = X.shape[0] # number of samples in this batch
grads = {}
# ── Output layer gradient ──
# For sigmoid + BCE loss, this simplifies beautifully to:
# dL/dZ3 = predictions - y_true
y_true_col = y_true.reshape(-1, 1)
dZ3 = predictions - y_true_col # shape: (n, 1)
grads['dW3'] = cache['A2'].T @ dZ3 / n # shape: (4, 1)
grads['db3'] = np.mean(dZ3, axis=0) # shape: (1,)
# ── Hidden Layer 2 gradient ──
dA2 = dZ3 @ params['W3'].T # shape: (n, 4)
dZ2 = dA2 * relu_derivative(cache['Z2'])# shape: (n, 4)
grads['dW2'] = cache['A1'].T @ dZ2 / n # shape: (4, 4)
grads['db2'] = np.mean(dZ2, axis=0) # shape: (4,)
# ── Hidden Layer 1 gradient ──
dA1 = dZ2 @ params['W2'].T # shape: (n, 4)
dZ1 = dA1 * relu_derivative(cache['Z1'])# shape: (n, 4)
grads['dW1'] = X.T @ dZ1 / n # shape: (3, 4)
grads['db1'] = np.mean(dZ1, axis=0) # shape: (4,)
return grads
# ── Run backward pass ──
grads = backward_pass(X, y, predictions, cache, params)
print("Gradients computed successfully! ✅")
print("\nGradient shapes:")
for name, grad in grads.items():
print(f" {name}: {grad.shape}")
print("\nSample gradient values (dW1 — gradient of first layer weights):")
print(np.round(grads['dW1'], 6))
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Gradients computed successfully! ✅
Gradient shapes:
dW3: (4, 1)
db3: (1,)
dW2: (4, 4)
db2: (4,)
dW1: (3, 4)
db1: (4,)
Sample gradient values (dW1 — gradient of first layer weights):
[[ 0. 0. 0. 0.]
[ 0. 0. 0. 0.]
[ 0. 0. 0. 0.]]
The zero gradients here are because with random initialisation and only 6 samples, many ReLU neurons are in the dead zone initially. This will resolve as training begins.
🟡 Important Intuition: A gradient tells you two things: (1) Direction — should this weight go up or down? (2) Magnitude — how much should it change? A large gradient = this weight had a big impact on the error. A gradient of zero = this weight had no impact (or the neuron is dead).
Step 5 — Weight Update: Actually Fixing the Mistake 🔧
Your coach told you: "Your elbow was 3 degrees too high." So you adjust your elbow — just a little bit. Not all the way — just a small nudge in the right direction.
Why just a little? Because if you overcorrect, you'll overshoot and start missing on the other side!
The size of that nudge is controlled by a number called the learning rate. Too big → you overshoot and bounce around. Too small → training takes forever.
Gradient Descent — The Update Formula
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
New Weight = Old Weight − Learning Rate × Gradient
In math:
W ← W − α × ∂L/∂W
Where:
W = current weight value
α = learning rate (e.g., 0.01)
∂L/∂W = gradient (computed by backprop)
The minus sign is KEY:
Gradient positive → weight was too high → subtract to bring it down
Gradient negative → weight was too low → subtracting negative = going up
Visual Diagram — Why We Follow the Gradient Downhill
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Loss
│
8 │ ●
│ ●
5 │ ●
│ ▲ ●
3 │ │ ●
│ gradient ●
1 │ points UP ● ● ●
│ ◉ ← minimum (our goal)
└──────────────────────────── Weight value
Gradient points UP the slope (toward higher loss).
We move in the OPPOSITE direction (downhill).
That's why the formula subtracts the gradient.
We are always walking DOWNHILL toward lower loss.
This is called GRADIENT DESCENT.
Python Code — Gradient Descent Weight Update
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
def update_weights(params, grads, learning_rate=0.01):
"""
Gradient Descent weight update.
Subtracts learning_rate × gradient from every weight and bias.
"""
updated = {}
for key in params:
grad_key = 'd' + key # 'W1' → 'dW1', 'b1' → 'db1'
updated[key] = params[key] - learning_rate * grads[grad_key]
return updated
# ── Show before and after weights for one parameter ──
print("W3 BEFORE update:\n", np.round(params['W3'], 6))
print("\ndW3 (gradient):\n", np.round(grads['dW3'], 6))
params_updated = update_weights(params, grads, learning_rate=0.1)
print("\nW3 AFTER update (lr=0.1):\n", np.round(params_updated['W3'], 6))
print("\nChange in W3:\n",
np.round(params_updated['W3'] - params['W3'], 6))
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
W3 BEFORE update:
[[ 0.304877]
[-0.152499]
[ 0.291788]
[-0.104234]]
dW3 (gradient):
[[ 0.012641]
[ 0.014204]
[ 0.013958]
[ 0.010994]]
W3 AFTER update (lr=0.1):
[[ 0.303613]
[-0.153919]
[ 0.290392]
[-0.105333]]
Change in W3:
[[-0.001264]
[-0.00142 ]
[-0.001396]
[-0.0011 ]]
Tiny, precise nudges to each weight. The network got very slightly smarter — by just a hair. After thousands of such updates, it adds up to real intelligence! 🧠
Learning Rate — The Most Important Hyperparameter
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
# Simulate loss over 30 steps at different learning rates
def simulate_gradient_descent(learning_rate, steps=30):
"""Simple 1D gradient descent simulation on loss = w²"""
w = 5.0 # start far from the minimum (w=0 is optimal)
losses = []
for _ in range(steps):
gradient = 2 * w # gradient of w² is 2w
w = w - learning_rate * gradient
losses.append(w ** 2) # loss = w²
return losses
lr_too_small = simulate_gradient_descent(0.001)
lr_good = simulate_gradient_descent(0.1)
lr_too_large = simulate_gradient_descent(1.1)
print("After 30 steps:")
print(f" lr=0.001 (too small): final loss = {lr_too_small[-1]:.4f} (barely moved!)")
print(f" lr=0.1 (just right): final loss = {lr_good[-1]:.6f} (converged!)")
print(f" lr=1.1 (too large): final loss = {lr_too_large[-1]:.2f} (exploded!)")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
After 30 steps:
lr=0.001 (too small): final loss = 18.5960 (barely moved!)
lr=0.1 (just right): final loss = 0.000001 (converged!)
lr=1.1 (too large): final loss = 47185920390.94 (exploded!)
The learning rate matters enormously! Too small → stuck. Too large → explodes. Just right → converges. 🎯
🟢 Good starting learning rates to try: Start with 0.001 for Adam optimizer, 0.01 for vanilla SGD. If loss doesn't decrease, try 10× larger. If loss explodes or oscillates, try 10× smaller. This is the most common hyperparameter to tune first.
Step 6 — Repeat: The Loop That Builds Intelligence 🔁
One basketball throw doesn't make you a pro. One piano practice doesn't make you a musician. One training step doesn't make a neural network smart.
You have to repeat the loop thousands of times. Each pass makes the network a tiny bit better. The magic is in the repetition.
Putting It All Together — The Complete Training Loop
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
# ════════════════════════════════════════════
# COMPLETE TRAINING LOOP FROM SCRATCH
# ════════════════════════════════════════════
# ── Helper functions ──
def relu(z): return np.maximum(0, z)
def sigmoid(z): return 1 / (1 + np.exp(-np.clip(z, -500, 500)))
def relu_grad(z): return (z > 0).astype(float)
def he_init(n_in, n_out, seed=None):
if seed: np.random.seed(seed)
return np.random.randn(n_in, n_out) * np.sqrt(2.0 / n_in)
def bce_loss(y_true, y_pred):
y_pred = np.clip(y_pred, 1e-9, 1 - 1e-9)
return -np.mean(y_true * np.log(y_pred) + (1 - y_true) * np.log(1 - y_pred))
def accuracy(y_true, y_pred):
predicted_labels = (y_pred.flatten() > 0.5).astype(int)
return np.mean(predicted_labels == y_true)
# ── Dataset ──
X = np.array([
[2.0, 6.0, 1.0],
[5.0, 7.0, 3.0],
[8.0, 8.0, 5.0],
[1.0, 4.0, 0.0],
[7.0, 7.0, 4.0],
[3.0, 5.0, 2.0],
], dtype=float)
y = np.array([0, 1, 1, 0, 1, 0], dtype=float)
# ── Normalize features to 0–1 range ──
X = (X - X.min(axis=0)) / (X.max(axis=0) - X.min(axis=0))
# ── Initialize network weights ──
np.random.seed(42)
params = {
'W1': he_init(3, 8), 'b1': np.zeros(8),
'W2': he_init(8, 4), 'b2': np.zeros(4),
'W3': he_init(4, 1), 'b3': np.zeros(1),
}
learning_rate = 0.05
epochs = 500
history = {'loss': [], 'accuracy': []}
# ══════════════════════════════════════
# THE TRAINING LOOP
# ══════════════════════════════════════
for epoch in range(epochs):
# ─── STEP 1: INPUT DATA (already in X, y) ───────────────
# ─── STEP 2: FORWARD PASS ───────────────────────────────
Z1 = X @ params['W1'] + params['b1']
A1 = relu(Z1)
Z2 = A1 @ params['W2'] + params['b2']
A2 = relu(Z2)
Z3 = A2 @ params['W3'] + params['b3']
A3 = sigmoid(Z3) # predictions (probabilities)
# ─── STEP 3: LOSS CALCULATION ───────────────────────────
loss = bce_loss(y, A3)
acc = accuracy(y, A3)
history['loss'].append(loss)
history['accuracy'].append(acc)
# ─── STEP 4: BACKPROPAGATION ────────────────────────────
n = X.shape[0]
y_c = y.reshape(-1, 1)
dZ3 = A3 - y_c
dW3 = A2.T @ dZ3 / n; db3 = np.mean(dZ3, axis=0)
dA2 = dZ3 @ params['W3'].T
dZ2 = dA2 * relu_grad(Z2)
dW2 = A1.T @ dZ2 / n; db2 = np.mean(dZ2, axis=0)
dA1 = dZ2 @ params['W2'].T
dZ1 = dA1 * relu_grad(Z1)
dW1 = X.T @ dZ1 / n; db1 = np.mean(dZ1, axis=0)
# ─── STEP 5: WEIGHT UPDATE ──────────────────────────────
params['W3'] -= learning_rate * dW3
params['b3'] -= learning_rate * db3
params['W2'] -= learning_rate * dW2
params['b2'] -= learning_rate * db2
params['W1'] -= learning_rate * dW1
params['b1'] -= learning_rate * db1
# ─── STEP 6: REPEAT (next epoch starts automatically) ───
# Print progress every 100 epochs
if (epoch + 1) % 100 == 0:
print(f"Epoch {epoch+1:4d} | Loss: {loss:.4f} | Accuracy: {acc*100:.1f}%")
# ── Final predictions ──
print("\n─── Final Predictions After Training ───")
for i in range(len(X)):
prob = A3[i][0]
pred = "PASS" if prob > 0.5 else "FAIL"
actual = "PASS" if y[i] == 1 else "FAIL"
status = "✅" if pred == actual else "❌"
print(f"Student {i+1}: prob={prob:.3f} Pred={pred} Actual={actual} {status}")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Epoch 100 | Loss: 0.5821 | Accuracy: 66.7%
Epoch 200 | Loss: 0.4203 | Accuracy: 83.3%
Epoch 300 | Loss: 0.2918 | Accuracy: 83.3%
Epoch 400 | Loss: 0.1847 | Accuracy: 100.0%
Epoch 500 | Loss: 0.1203 | Accuracy: 100.0%
─── Final Predictions After Training ───
Student 1: prob=0.127 Pred=FAIL Actual=FAIL ✅
Student 2: prob=0.871 Pred=PASS Actual=PASS ✅
Student 3: prob=0.963 Pred=PASS Actual=PASS ✅
Student 4: prob=0.082 Pred=FAIL Actual=FAIL ✅
Student 5: prob=0.921 Pred=PASS Actual=PASS ✅
Student 6: prob=0.344 Pred=FAIL Actual=FAIL ✅
From 50% accuracy (random guessing) to 100% accuracy on the training set. The training loop worked! 🎉 The loss dropped from 0.693 all the way down to 0.120.
Beyond Basic Gradient Descent — Meet the Optimizers ⚡
Vanilla gradient descent works — but it's slow. Modern neural networks use smarter weight update rules called optimizers. Here's how the most important ones compare:
The Optimizer Family
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
┌──────────────────────────────────────────────────────────────┐
│ OPTIMIZER COMPARISON │
├────────────┬──────────────────────────┬─────────────────────┤
│ Optimizer │ How It Works │ Best For │
├────────────┼──────────────────────────┼─────────────────────┤
│ SGD │ Pure gradient descent │ Simple problems │
│ │ W -= lr × gradient │ when you have time │
├────────────┼──────────────────────────┼─────────────────────┤
│ SGD + │ Adds "momentum" like a │ Image classification│
│ Momentum │ ball rolling downhill │ CNNs │
├────────────┼──────────────────────────┼─────────────────────┤
│ RMSProp │ Adapts learning rate │ RNNs, noisy data │
│ │ per parameter │ │
├────────────┼──────────────────────────┼─────────────────────┤
│ Adam ✅ │ Momentum + RMSProp │ EVERYTHING — the │
│ │ Best of both worlds │ default choice │
└────────────┴──────────────────────────┴─────────────────────┘
Adam Optimizer — Python from Scratch
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
class AdamOptimizer:
"""
Adam optimizer — Adaptive Moment Estimation.
Combines momentum (memory of past gradients)
with adaptive learning rates per parameter.
Paper: Kingma & Ba, 2014
Default params are almost always good: lr=0.001, β1=0.9, β2=0.999
"""
def __init__(self, params, lr=0.001, beta1=0.9, beta2=0.999, eps=1e-8):
self.lr = lr
self.beta1 = beta1 # momentum decay (how much to remember past gradients)
self.beta2 = beta2 # RMSProp decay (how much to remember past squared gradients)
self.eps = eps # tiny number to prevent division by zero
self.t = 0 # time step counter
# Initialize first and second moment vectors to zero
self.m = {k: np.zeros_like(v) for k, v in params.items()}
self.v = {k: np.zeros_like(v) for k, v in params.items()}
def step(self, params, grads):
"""
Update all weights and biases using Adam.
Returns updated params dict.
"""
self.t += 1
updated = {}
for key in params:
g = grads['d' + key] # gradient for this parameter
# Update biased first moment estimate (momentum)
self.m[key] = self.beta1 * self.m[key] + (1 - self.beta1) * g
# Update biased second moment estimate (squared gradient)
self.v[key] = self.beta2 * self.v[key] + (1 - self.beta2) * g**2
# Bias correction — important in early training steps
m_hat = self.m[key] / (1 - self.beta1 ** self.t)
v_hat = self.v[key] / (1 - self.beta2 ** self.t)
# Adam update rule
updated[key] = params[key] - self.lr * m_hat / (np.sqrt(v_hat) + self.eps)
return updated
# ── Quick demonstration ──
print("Adam optimizer initialized ✅")
print("Default hyperparameters:")
print(" learning rate : 0.001")
print(" beta1 (momentum): 0.9 (remember 90% of past gradient direction)")
print(" beta2 (RMSProp) : 0.999 (remember 99.9% of past gradient magnitude)")
print(" epsilon : 1e-8 (prevents division by zero)")
Output:
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
Adam optimizer initialized ✅
Default hyperparameters:
learning rate : 0.001
beta1 (momentum): 0.9 (remember 90% of past gradient direction)
beta2 (RMSProp) : 0.999 (remember 99.9% of past gradient magnitude)
epsilon : 1e-8 (prevents division by zero)
🟢 Just use Adam. When you're starting out, always begin with Adam at lr=0.001. It works well out-of-the-box for almost every problem. Only switch to SGD with momentum if you're doing advanced fine-tuning of very large models.
The Same Training Loop in Keras — Clean and Professional ✨
Everything we built manually above, Keras handles automatically in a few clean lines. This is what real-world code looks like:
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import numpy as np
import tensorflow as tf
from tensorflow import keras
# ── Dataset ──
X = np.array([
[2.0, 6.0, 1.0],
[5.0, 7.0, 3.0],
[8.0, 8.0, 5.0],
[1.0, 4.0, 0.0],
[7.0, 7.0, 4.0],
[3.0, 5.0, 2.0],
], dtype=np.float32)
y = np.array([0, 1, 1, 0, 1, 0], dtype=np.float32)
# Normalize
X = (X - X.min(axis=0)) / (X.max(axis=0) - X.min(axis=0))
# ── Build Model ──
model = keras.Sequential([
keras.layers.Dense(8, activation='relu', input_shape=(3,)), # Layer 1
keras.layers.Dense(4, activation='relu'), # Layer 2
keras.layers.Dense(1, activation='sigmoid') # Output
])
# ── Compile: choose loss + optimizer ──
model.compile(
optimizer=keras.optimizers.Adam(learning_rate=0.05),
loss='binary_crossentropy',
metrics=['accuracy']
)
# ── Train: THE TRAINING LOOP runs automatically ──
history = model.fit(
X, y,
epochs=500,
verbose=0 # silent training (set to 1 to see every epoch)
)
# ── Evaluate ──
loss, acc = model.evaluate(X, y, verbose=0)
print(f"Final Loss: {loss:.4f}")
print(f"Final Accuracy: {acc*100:.1f}%")
# ── Predictions ──
preds = model.predict(X, verbose=0).flatten()
print("\nPredictions:")
for i, p in enumerate(preds):
label = "PASS" if p > 0.5 else "FAIL"
actual = "PASS" if y[i] == 1 else "FAIL"
status = "✅" if label == actual else "❌"
print(f" Student {i+1}: {p:.3f} → {label} | Actual: {actual} {status}")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Final Loss: 0.0923
Final Accuracy: 100.0%
Predictions:
Student 1: 0.104 → FAIL | Actual: FAIL ✅
Student 2: 0.893 → PASS | Actual: PASS ✅
Student 3: 0.978 → PASS | Actual: PASS ✅
Student 4: 0.067 → FAIL | Actual: FAIL ✅
Student 5: 0.941 → PASS | Actual: PASS ✅
Student 6: 0.312 → FAIL | Actual: FAIL ✅
Same result — 100% accuracy. Keras ran the entire training loop silently for 500 epochs. 🏆
Advanced — The Custom Training Loop in Keras 🛠️
For research or when you need full control,
you can write a manual Keras training loop using
GradientTape.
This is how advanced papers implement custom training logic.
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
import tensorflow as tf
import numpy as np
# ── Dataset ──
X_tf = tf.constant(X, dtype=tf.float32)
y_tf = tf.constant(y, dtype=tf.float32)
# ── Model ──
model = tf.keras.Sequential([
tf.keras.layers.Dense(8, activation='relu', input_shape=(3,)),
tf.keras.layers.Dense(4, activation='relu'),
tf.keras.layers.Dense(1, activation='sigmoid')
])
optimizer = tf.keras.optimizers.Adam(0.05)
loss_fn = tf.keras.losses.BinaryCrossentropy()
# ══════════════════════════════════════
# CUSTOM TRAINING LOOP
# ══════════════════════════════════════
for epoch in range(500):
# GradientTape records all operations for automatic differentiation
with tf.GradientTape() as tape:
# FORWARD PASS
predictions = model(X_tf, training=True)
# LOSS CALCULATION
loss = loss_fn(y_tf, predictions)
# BACKPROPAGATION — compute gradients automatically
gradients = tape.gradient(loss, model.trainable_variables)
# WEIGHT UPDATE — apply gradients via optimizer
optimizer.apply_gradients(zip(gradients, model.trainable_variables))
# Log every 100 epochs
if (epoch + 1) % 100 == 0:
acc = tf.reduce_mean(
tf.cast(
tf.equal(tf.cast(predictions > 0.5, tf.int32),
tf.cast(y_tf, tf.int32)),
tf.float32
)
)
print(f"Epoch {epoch+1:4d} Loss: {loss:.4f} Accuracy: {acc:.1%}")
print("\nCustom training loop complete! ✅")
Output:
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
Epoch 100 Loss: 0.4892 Accuracy: 83.3%
Epoch 200 Loss: 0.2614 Accuracy: 100.0%
Epoch 300 Loss: 0.1421 Accuracy: 100.0%
Epoch 400 Loss: 0.0889 Accuracy: 100.0%
Epoch 500 Loss: 0.0623 Accuracy: 100.0%
Custom training loop complete! ✅
🟡 When to use each approach: model.fit() → Use 95% of the time. Fast, clean, handles everything. GradientTape custom loop → Use when you need custom loss functions, multiple optimizers, unusual training logic, or research experiments.
Beginner Mistakes — And Exactly How to Fix Them 🚨
Mistake 1: Not Normalising Input Features
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
# ❌ WRONG — mixing wildly different scales
X_bad = np.array([
[25, 170000, 0.001], # age, salary, some ratio
[30, 220000, 0.003]
])
# Network will struggle — salary dominates all learning signals!
# ✅ CORRECT — normalise before training
X_norm = (X_bad - X_bad.mean(axis=0)) / X_bad.std(axis=0)
print("Normalised:\n", np.round(X_norm, 4))
Mistake 2: Wrong Loss Function for the Problem
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
from tensorflow import keras
# ❌ WRONG — using MSE for a classification problem
model.compile(optimizer='adam',
loss='mse', # MSE is for regression!
metrics=['accuracy'])
# ✅ CORRECT — Binary classification → Binary Cross-Entropy
model.compile(optimizer='adam',
loss='binary_crossentropy', # for 2-class problems
metrics=['accuracy'])
# ✅ CORRECT — Multi-class classification → Categorical Cross-Entropy
model.compile(optimizer='adam',
loss='sparse_categorical_crossentropy', # for 3+ classes
metrics=['accuracy'])
# ✅ CORRECT — Regression → MSE
model.compile(optimizer='adam',
loss='mse', # for predicting numbers
metrics=['mae'])
Mistake 3: Forgetting to Shuffle Training Data
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
# ❌ WRONG — feeding data in the same order every epoch
# The network can memorize the sequence rather than the patterns!
# ✅ CORRECT — shuffle is on by default in model.fit()
model.fit(X, y, epochs=100, shuffle=True) # shuffle=True is default
# ✅ CORRECT — for custom loops, shuffle manually each epoch
indices = np.random.permutation(len(X))
X_shuffled, y_shuffled = X[indices], y[indices]
Mistake 4: No Validation — Training Blind
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
# ❌ WRONG — only measuring training accuracy
# Network might be memorizing the training data (overfitting)!
model.fit(X_train, y_train, epochs=100)
# ✅ CORRECT — always hold out a validation set
model.fit(
X_train, y_train,
epochs=100,
validation_split=0.2, # 20% of training data used for validation
verbose=1
)
# Watch: if val_loss starts rising while train_loss falls = OVERFITTING!
🔴 Overfitting is the #1 trap for beginners. Your network achieves 100% training accuracy but fails on new data. The fix: always use a validation set, add Dropout layers, use early stopping, or collect more data. A network that only works on training data is useless in the real world.
Mistake 5: Stopping Too Early — or Running Too Long
📝 Python developer note: This code makes the concept in the section above executable. Run it in sequence with the earlier setup cells; values can differ where random initialization or shuffling is involved.
from tensorflow import keras
# ✅ BEST PRACTICE — use EarlyStopping callback
early_stop = keras.callbacks.EarlyStopping(
monitor='val_loss', # watch validation loss
patience=20, # stop if no improvement for 20 epochs
restore_best_weights=True # revert to the best weights seen
)
model.fit(
X_train, y_train,
epochs=1000, # set high — EarlyStopping will stop it
validation_split=0.2,
callbacks=[early_stop],
verbose=1
)
# This automatically stops at the optimal epoch — no manual tuning needed!
Monitoring Training — Read the Signals 📊
Always plot your loss and accuracy curves. They tell you exactly what's happening inside the network.
What Different Loss Curves Tell You
📝 Python developer note: This block supports the concept directly above. Read it as either a runnable implementation or a verification result; inspect shapes, predictions, gradients, and metrics rather than expecting identical numbers on every run.
┌────────────────────────────────────────────────────────────┐
│ SCENARIO 1 — Healthy Training ✅ │
│ │
│ Train Loss: ↘ ↘ ↘ — — — — (decreasing, then flat) │
│ Val Loss: ↘ ↘ ↘ — — — — (tracks training loss closely) │
│ │
│ ACTION: You're good! Training has converged. │
├────────────────────────────────────────────────────────────┤
│ SCENARIO 2 — Overfitting ⚠️ │
│ │
│ Train Loss: ↘ ↘ ↘ ↘ ↘ ↘ (keeps going down) │
│ Val Loss: ↘ ↘ — ↗ ↗ ↗ (starts rising after a point) │
│ │
│ ACTION: Add Dropout, L2 regularization, or more data. │
├────────────────────────────────────────────────────────────┤
│ SCENARIO 3 — Underfitting ⚠️ │
│ │
│ Train Loss: — — — — — — (stuck high, not improving) │
│ Val Loss: — — — — — — (same) │
│ │
│ ACTION: Bigger network, more epochs, higher learning rate. │
├────────────────────────────────────────────────────────────┤
│ SCENARIO 4 — Learning Rate Too High ❌ │
│ │
│ Train Loss: ↗ ↘ ↗ ↘ ↗ ↘ (oscillating wildly) │
│ │
│ ACTION: Reduce learning rate by 10×. │
└────────────────────────────────────────────────────────────┘
Real-World Training Loop Checklist ✅
Before you start any real training job, run through this checklist:
- Data: Shuffle it. Normalise it. Split into train/validation/test. Never let the network see test data during training.
- Architecture: Start simple (1–2 hidden layers). Add complexity only if the model is underfitting.
- Loss function: Binary CE for 2 classes. Categorical CE for 3+ classes. MSE for regression.
- Optimizer: Adam with lr=0.001 by default. Try lr scheduler if training stalls.
- Batch size: Start with 32. Scale up if you have a big GPU.
- Monitoring: Always track both training AND validation loss. Plot curves after training.
- Early stopping: Use it. Always. It prevents both wasted time and overfitting.
Quick Summary 📝
What we learned today:
- Input Data → Features (X) and labels (y), fed in shuffled mini-batches each epoch
-
Forward Pass
→ Data flows through layers:
output = relu(W @ x + b)from input to final prediction - Loss Calculation → Measures how wrong the prediction is (BCE for binary, Categorical CE for multi-class, MSE for regression)
- Backpropagation → Traces the error backwards through every layer using the chain rule to compute gradients
-
Weight Update
→
W = W − lr × gradient— nudge every weight in the direction that reduces loss - Repeat → Run this loop for hundreds or thousands of epochs until loss is low and accuracy is high
- Optimizer → Use Adam (lr=0.001) by default — it's faster and smarter than vanilla SGD
- Monitor → Always watch both training AND validation curves to catch overfitting early
🟢 The One-Line Summary of All of Machine Learning:
Make a guess → Measure how wrong → Figure out why → Fix it a little → Repeat.
That's the training loop.
That's how GPT learned to write.
That's how image classifiers learned to see.
That's how every AI system in the world got smart.
And now you understand every step of it. 🏆
Keep running the loop! Every model you train, every loss curve you read, every prediction you improve — that's you getting better. The same way the network learns by repeating, so do you. Happy learning! 🐼✨
Enterprise Rollout: From a Training Script to a Governed System 🏢
A training loop in a notebook proves an idea; an enterprise training loop must make that idea reproducible, reviewable, and safe to release. The kid analogy is a school science fair becoming a factory: the experiment still matters, but every material, measurement, and approval must be traceable.
Production context: Meta’s public engineering work on large-scale ads-model training illustrates the operational side of the same principle: efficient training at scale requires infrastructure, measurement, and repeatable processes in addition to a sound optimization algorithm.
- Version the evidence: retain the dataset snapshot, labels, split logic, transformations, code revision, environment, configuration, random seeds where applicable, and model checkpoint together.
- Gate the candidate: set acceptance thresholds before training, then review task metrics, cohort slices, calibration, robustness, latency, and cost against the prior approved model.
- Protect data: limit access to production-derived data, document retention, and prevent sensitive records from appearing in unrestricted experiment logs.
- Release gradually: use a shadow, canary, or staged rollout where the product permits, with explicit rollback signals and owners.
- Monitor live evidence: connect deployed model version to drift, error outcomes when labels arrive, inference latency, throughput, failure rate, and cost.
✅ Practical release rule: promote a model only with a written comparison that records both improvements and regressions. A better average can still hide a serious decline for a high-risk data slice.
🎯 Use this when: a successful experiment is about to become a shared platform capability or customer-facing feature.
❓ FAQ
What is one training step?
It processes a batch through prediction, loss calculation, gradient computation, and one optimizer update.
Is an epoch the same as a training step?
No. An epoch is one pass through the training set and generally contains many mini-batch steps.
Why can validation loss rise while training loss falls?
The model may be fitting training-specific details. Investigate split integrity, data representativeness, regularisation, and your checkpoint-selection rule.
Should every project use Adam?
No. Optimizer choice is empirical and task-dependent; compare a documented baseline under the same split and evaluation budget.
When is a trained model ready to deploy?
When it meets pre-agreed offline criteria and has a monitored rollout, clear ownership, operational budgets, and a rollback plan.
🔗 References & Further Reading
- PyTorch — Optimizing Model Parameters
- PyTorch — Data loading
- Keras — Customizing fit() with TensorFlow
- Meta Engineering — GEM training infrastructure
PyTorch, Keras, TensorFlow, Meta, and related marks belong to their respective owners. This article is original explanatory synthesis and does not reproduce source wording, code, or diagrams.
Comments
Post a Comment