Learning rate and batch size are the two dials that control how a model's training loop actually moves through its search for good weights: the learning rate sets how big each step is, and the batch size sets how many examples inform each step before it's taken. Schedulers and warm-up are refinements on top of the learning rate dial — instead of picking one fixed value, they change it deliberately over the course of training. Get these wrong and even a perfectly correct model architecture will train slowly, unstably, or not at all. 🎛️
Of all the hyperparameters in deep learning, these four are usually the first ones worth tuning, because they affect every single training run regardless of architecture, and because the failure modes when they're wrong are often confusing — a loss that explodes to NaN, a model that seems stuck, or a large-batch run that mysteriously underperforms a small-batch one on the exact same data. This post builds up the reasoning from first principles, then shows the real mechanisms frameworks and published research use to handle them. 📈
📑 In This Post
- 1. Foundations: The Two Dials That Control Training
- 2. Learning Rate: Too High, Too Low, Just Right
- 3. Batch Size: What It Actually Changes
- 4. The Learning Rate–Batch Size Relationship: Linear Scaling
- 5. Warm-Up: Why Start Slow
- 6. Schedulers: Changing the Learning Rate Over Time
- 7. Implementation: Warm-Up Plus Decay in Code
- 8. Enterprise Rollout: Governing Hyperparameters at Scale
- 9. Common Mistakes
- 10. FAQ
- 11. References & Further Reading
- 12. Summary
🔀 Quick Comparison: Small Batch Size vs. Large Batch Size
| Property | Small batch size | Large batch size |
|---|---|---|
| Gradient estimate | Noisier — based on fewer examples | Smoother — closer to the true average gradient |
| Hardware efficiency | Underuses parallel hardware per step | Better utilizes GPU/TPU parallelism per step |
| Steps per epoch | More steps for the same data | Fewer steps for the same data |
| Learning rate needed | Typically smaller | Typically larger, scaled up with care |
| Memory use | Lower | Higher — more activations held at once |
1. Foundations: The Two Dials That Control Training
💭 Analogy first: imagine hiking down a foggy hill to find the lowest point in the valley, using only your feet to feel which way is downhill. The learning rate is your stride length — too short and you'll take forever to get anywhere; too long and you might stride right over the valley floor and end up higher on the opposite slope. The batch size is how many nearby spots you check with your foot before deciding which direction "downhill" really is — checking more spots gives a more reliable direction, but takes longer per step.
Every training step nudges the model's weights in the direction that (locally) reduces the loss. The learning rate is the size of that nudge. The batch size is how many training examples are averaged together to decide which direction the nudge should go. Together they determine how training actually behaves: how fast it moves, how noisy or smooth its path is, and how efficiently it uses the hardware underneath it.
Schedulers and warm-up build on top of the learning rate dial specifically. Rather than fixing one learning rate for the entire run, a scheduler changes it according to a plan — usually starting small (warm-up), rising, and then decaying — because the "right" step size early in training is often not the right step size once the model is close to a good solution.
🎯 Use this when you need the one-paragraph mental model before diving into any specific number or formula.
2. Learning Rate: Too High, Too Low, Just Right
💭 Analogy first: setting an oven too hot burns the outside of a cake before the inside cooks — the process overshoots before it can finish properly. Set it too low, and the cake technically cooks eventually, but far more slowly than necessary, and you might give up and pull it out too early, assuming it's simply not going to work.
What happens when the learning rate is too high: each step overshoots the direction it should have moved in. Instead of gradually settling toward a good set of weights, the loss oscillates, and in the worst case grows without bound — visible as a loss curve that spikes upward or turns into NaN (not-a-number) entirely.
What happens when it's too low: every step is technically in the right direction, but so tiny that visible progress can take an impractically long time, or the model gets stuck in a shallow, mediocre region of the loss landscape it could have escaped with slightly bigger steps.
What "just right" looks like: a loss curve that decreases steadily without wild oscillation, ideally with a shape that flattens out gradually as the model approaches a good solution rather than plateauing very early or diverging.
💡 Trade-off: there is no single universally correct learning rate — it depends on the model architecture, the optimizer, the batch size, and the data. This is exactly why schedulers and warm-up exist: instead of finding one perfect fixed number, they let the "right" value change as training progresses.
3. Batch Size: What It Actually Changes
💭 Analogy first: polling 10 random people about an election gives you a rough, noisy sense of public opinion; polling 10,000 gives a much more stable, reliable estimate, but takes proportionally more time and effort to collect. A training batch is a poll of the dataset's gradient direction.
What it does: a batch is the group of examples used together to compute one gradient estimate before a weight update. A larger batch averages over more examples, producing a gradient estimate that's closer to what you'd get from the entire dataset (the "true" gradient); a smaller batch produces a noisier estimate based on fewer examples.
Why it matters beyond noise: hardware like GPUs and TPUs is built for parallel computation, so processing a larger batch in one step is usually far more efficient per-example than processing many small batches sequentially — up to the point where the batch no longer fits in accelerator memory.
What fails at the extremes: an extremely small batch size (like 1) produces such a noisy gradient estimate that training can wander unpredictably and slowly; an extremely large batch size, used naively with the same learning rate as a small batch, tends to make less progress per epoch than expected, because fewer, smoother steps are taken for the same amount of data seen — which is precisely the problem the linear scaling rule in the next section addresses.
4. The Learning Rate–Batch Size Relationship: Linear Scaling
This is one of the most concretely documented relationships in the field, thanks to a well-known 2017 Facebook AI Research paper on training ResNet-50 with very large batches. The paper's own stated rule: when the minibatch size is multiplied by k, multiply the learning rate by k as well, while keeping all other hyperparameters, such as weight decay, unchanged.
Why this rule makes intuitive sense: a larger batch produces a smoother, more averaged gradient direction, but by itself that doesn't mean a bigger step is safe — the paper's own reasoning is that k small-batch steps at learning rate η should behave similarly to one large-batch step of k times the size, only if the large step's learning rate is also scaled up by k, so the total movement per amount-of-data-seen stays comparable.
Real, documented result. Using this linear scaling rule together with a warm-up strategy (covered next), the paper's authors trained a ResNet-50 model on ImageNet with a minibatch size of 8,192 across 256 GPUs in one hour, while matching the accuracy of the standard, much smaller 256-image minibatch baseline — and reported roughly 90% scaling efficiency when moving from 8 to 256 GPUs.
✅ Worked example: if a baseline recipe uses batch size 256 with learning rate 0.1, the linear scaling rule suggests batch size 8,192 (32 times larger) should pair with learning rate 3.2 (32 times larger) — not the same 0.1, which the documented research shows would under-train the larger-batch run relative to its potential.
The important caveat, from the same research: the paper explicitly notes this rule alone is not sufficient — a large, suddenly-applied learning rate causes early-training instability, which is exactly why the same paper introduces a warm-up scheme rather than just scaling the learning rate and starting training normally.
5. Warm-Up: Why Start Slow
💭 Analogy first: a car engine on a cold morning runs roughly if you immediately floor the accelerator, but drives smoothly once it's warmed up for a minute first. Warm-up gives a freshly initialized model's very first updates a gentler ride before letting it move at full speed.
What it does: instead of starting training at the full target learning rate, warm-up begins at a very small (sometimes zero) learning rate and increases it, usually linearly, over a fixed number of initial steps before switching to the main schedule.
Why it's needed: at the very start of training, weights are randomly initialized and the optimizer's internal statistics (for adaptive optimizers like Adam) haven't stabilized yet, so a large learning rate applied immediately can push weights into a bad region or cause the loss to spike before the model has had any chance to find a reasonable starting direction.
Real, documented example — the original Transformer. The "Attention Is All You Need" paper defines its learning rate explicitly as a function of the training step number: it increases the learning rate linearly for the first warmup_steps training steps, then decreases it proportionally to the inverse square root of the step number afterward. The paper states it used warmup_steps = 4000, alongside the Adam optimizer with β1 = 0.9, β2 = 0.98, and ε = 10⁻⁹.
Real, documented example — large-batch image classification. The same large-minibatch ResNet-50 research referenced above developed its own warm-up scheme specifically to make the linear-scaling learning rate usable, describing it as overcoming "optimization challenges early in training" that appear when a large learning rate is applied to a freshly initialized network without any ramp-up period.
What fails without warm-up when using a large or scaled-up learning rate: both cited papers point to the same failure mode — instability and degraded final accuracy in the first phase of training — which is precisely why warm-up and large-batch linear scaling are documented together rather than as separate, unrelated techniques.
6. Schedulers: Changing the Learning Rate Over Time
A scheduler is the general mechanism for changing the learning rate according to a plan across training, and warm-up is simply the first phase of many real-world schedules. PyTorch's torch.optim.lr_scheduler module documents a wide catalog of these strategies, several of which are worth knowing by name:
- StepLR / MultiStepLR: multiply the learning rate by a fixed factor (
gamma) every fixed number of epochs, or at specific named milestones. Simple and predictable, but the drops are abrupt. - ExponentialLR: multiply the learning rate by a constant factor every single epoch, producing smooth, continuous decay rather than sudden drops.
- CosineAnnealingLR: decay the learning rate following a smooth cosine curve from an initial value down toward a minimum, which tends to decelerate the rate of decay gently near the end rather than dropping sharply.
- OneCycleLR: combine a warm-up rise to a peak learning rate with a subsequent decay, both within a single specified total number of steps — effectively packaging warm-up and decay into one documented scheduler object.
- ReduceLROnPlateau: instead of following a fixed step-based plan, this scheduler watches a monitored metric (commonly validation loss) and reduces the learning rate only when that metric stops improving for a specified number of epochs.
An important documented operational detail: PyTorch's own source explicitly warns that as of PyTorch 1.1.0 and later, optimizer.step() must be called before scheduler.step() — calling them in the opposite order causes PyTorch to skip the first value of the learning rate schedule entirely, a documented gotcha that's easy to introduce when refactoring a training loop.
🎯 Use this section as a lookup table when you're not sure which named scheduler matches the behavior you want.
7. Implementation: Warm-Up Plus Decay in Code
Combining a linear warm-up with a cosine decay afterward is one of the most common real-world schedules. Here's an original, illustrative implementation using PyTorch's documented SequentialLR to chain two schedulers together.
Notice the comment marking the order of optimizer.step() and scheduler.step() — this is exactly the documented ordering requirement from Section 6, and getting it backward is one of the most common silent mistakes in real training scripts.
8. Enterprise Rollout: Governing Hyperparameters at Scale
Learning rate, batch size, and schedule choices aren't "set once" decisions on a production team — they're hyperparameters that need the same rigor as any other configuration that materially affects model quality and cost.
Version and log every hyperparameter with every run: learning rate, batch size, warm-up length, and scheduler type should be recorded alongside the resulting model checkpoint. Without this, comparing "why did last month's run do better" across models becomes guesswork.
Treat batch size changes as a re-tuning event, not a free lunch: because of the documented linear scaling relationship, increasing batch size to use more GPUs without revisiting the learning rate and warm-up length is a common way accuracy quietly regresses after an infrastructure change.
Divergence monitoring: alert automatically on a loss that spikes, or turns into NaN or Inf, so a run using a too-aggressive learning rate is caught and stopped within minutes rather than being discovered hours later when someone checks in on it.
Hyperparameter sweeps as a controlled, budgeted process: searching over learning rate and batch size combinations can consume enormous compute if unbounded; define a fixed sweep budget and a clear metric for choosing a winner in advance, rather than sweeping indefinitely.
Rollback criteria: if a new default learning rate schedule is rolled out across a team's training pipelines, define in advance what validation metric regression would trigger reverting to the previous default, and test the new schedule on a canary subset of jobs first.
✅ Practical pattern: log the learning rate value itself (not just the loss) on every training dashboard. A schedule that isn't behaving as configured — wrong warm-up length, an accidental order-of-operations bug — is usually obvious the moment you actually look at the logged learning rate curve.
9. Common Mistakes
Calling scheduler.step() before optimizer.step(). As PyTorch's own documentation warns, this causes the first scheduled learning rate value to be silently skipped — the training loop runs without any error, but the schedule is subtly off from what was intended for the entire run.
Scaling batch size without scaling learning rate. Moving from one GPU to eight and multiplying batch size accordingly, while leaving the learning rate untouched, routinely under-trains the model relative to what the documented linear scaling rule would predict — the run "works" but underperforms, without any obvious sign of what went wrong.
Skipping warm-up when scaling up the learning rate. Applying a large, linearly-scaled learning rate immediately at step one, without any ramp-up period, is exactly the failure mode the cited large-batch research and the Transformer paper both address directly — expect instability or a stalled early loss curve.
Restarting training from a checkpoint without restoring scheduler state. If only model and optimizer state are saved and reloaded but the scheduler is reinitialized from scratch, the learning rate silently jumps back to an earlier point in the schedule instead of continuing from where training left off.
Setting ReduceLROnPlateau's patience too short. Validation metrics naturally fluctuate a little between epochs; a patience window that's too aggressive can trigger learning rate drops on ordinary noise rather than genuine plateaus, causing the learning rate to decay far earlier than intended.
❓ FAQ
Should I always increase the learning rate when I increase batch size?
Generally yes, following something like the documented linear scaling rule as a starting point, but it should be verified rather than assumed — the same research that introduced the rule also found it needs to be paired with warm-up to work reliably, and very large batch sizes can eventually reach limits where linear scaling alone isn't sufficient.
How long should warm-up last?
There's no universal number — the original Transformer paper used 4,000 steps for its specific setup — and the right length depends on model size, learning rate magnitude, and batch size. As a starting point, a few percent of total training steps is a common practical choice, tuned based on whether early-training loss looks stable.
Which scheduler should a beginner start with?
A simple linear warm-up followed by cosine decay is a strong, widely used default for training from scratch. ReduceLROnPlateau is a reasonable alternative when you'd rather react to actual validation performance than commit to a fixed step-based plan in advance.
My loss became NaN early in training — is this a learning rate problem?
It's one of the most common causes, especially if the learning rate is high relative to the batch size or if warm-up is missing entirely. Try adding or lengthening warm-up and reducing the peak learning rate before investigating other causes like data issues or numerical instability elsewhere in the model.
Does a bigger batch size always mean faster training?
Not automatically. It can better utilize parallel hardware per step, but if the learning rate isn't scaled appropriately, fewer, larger steps can mean less effective progress per epoch, offsetting or even eliminating the wall-clock speed advantage.
🔗 References & Further Reading
- Goyal et al., "Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour" (2017) — the linear scaling rule and the large-batch warm-up scheme, with the ResNet-50/8,192-batch/256-GPU result.
- Vaswani et al., "Attention Is All You Need" (2017) — the documented Transformer learning rate schedule, Adam hyperparameters, and warmup_steps = 4000.
- PyTorch: torch.optim documentation — the catalog of learning rate schedulers, including StepLR, CosineAnnealingLR, OneCycleLR, ReduceLROnPlateau, and SequentialLR, and the optimizer.step()-before-scheduler.step() ordering requirement.
PyTorch is a trademark of its respective owners;
📝 Summary
- Learning rate sets the size of each training step; batch size sets how many examples inform the direction of that step.
- A learning rate that's too high causes instability or divergence; too low causes impractically slow progress.
- Larger batches give smoother gradient estimates and better hardware utilization, but need a correspondingly larger learning rate to make full use of that averaging.
- The documented linear scaling rule says: multiply the learning rate by the same factor you multiply the batch size by.
- Warm-up starts training at a small learning rate and ramps up, preventing instability from a large step applied to a freshly initialized model.
- Schedulers like StepLR, CosineAnnealingLR, OneCycleLR, and ReduceLROnPlateau each implement a different documented strategy for changing the learning rate over training.
- Calling optimizer.step() before scheduler.step() is a documented, easy-to-miss ordering requirement.
- Enterprise training pipelines need to log, version, and monitor these hyperparameters just as carefully as the model architecture itself.
Thanks for reading — may your loss curves warm up gently and decay smoothly. 🚀
Comments
Post a Comment