Skip to main content

What's the Difference Between SGD, Momentum, Adam, and AdamW

Calculating read time…

An optimizer is the algorithm that decides exactly how much to adjust every weight in a network after each training step, using the gradients that backpropagation just computed. SGD, Momentum, Adam, and AdamW are four steps along the same evolutionary line: each one keeps what worked in the last and fixes one specific, well-documented weakness. Understanding what each one actually changes about the update rule is the difference between guessing at hyperparameters and knowing exactly why a training run is or isn't converging. 🎯

This matters in production because the optimizer choice interacts directly with learning rate, weight decay, batch size, and training stability — get it wrong and you'll either waste enormous compute budget on a run that converges too slowly, or watch a seemingly healthy loss curve diverge without warning. Engineers who understand the actual update rule behind their optimizer's name can read a divergence or plateau and know immediately which knob to turn. ⚙️

A diagram of elliptical loss-surface contour lines representing a narrow valley, with two paths from the same starting point to the minimum: a jagged zig-zagging path labeled plain SGD that bounces back and forth across the valley walls, and a smoother, more direct path labeled SGD with momentum that curves steadily toward the minimum.

Figure 1. Original diagram: momentum turns a bouncing path into a smoother, more direct one.

🔀 Quick Comparison: SGD, Momentum, Adam, and AdamW

Property Plain SGD SGD + Momentum Adam AdamW
Per-parameter learning rate No No Yes, adaptive Yes, adaptive
Remembers past gradients No Yes, one running average Yes, two running averages Yes, two running averages
Weight decay handling N/A by default Added into the gradient (L2-style) Added into the gradient, then scaled adaptively Applied directly to weights, decoupled
Introduced Classic algorithm Classic algorithm Kingma & Ba, 2014 Loshchilov & Hutter, 2017–2019

1. Foundations: What an Optimizer's Update Rule Actually Is

🧠 Child-friendly analogy first: imagine you're walking down a foggy hill in the dark, feeling only the slope right under your feet at each step. Each of these optimizers is a different strategy for deciding how big a step to take and in which direction, using only that local slope information — some walk cautiously and forget every previous step, some build up speed by remembering which way they've been leaning, and some adjust their stride length separately for each leg.

Technically, every one of these optimizers takes the same raw ingredient — the gradient of the loss with respect to each parameter, computed by backpropagation — and turns it into a parameter update using its own specific formula. The differences between SGD, Momentum, Adam, and AdamW are entirely in that formula: what state each one remembers between steps, and how it uses that state to scale or redirect the current gradient before applying it.

Because all four optimizers consume the exact same gradients, swapping one for another never changes what backpropagation computes — it only changes how those gradients get turned into a weight update, which is why an optimizer swap alone can dramatically change training speed and stability without touching the model architecture at all.

🎯 Use this when you need a clear mental separation between "what backpropagation computes" and "what the optimizer does with it."

2. Plain SGD: The Baseline Update Rule

🧠 Analogy: think of someone with severe short-term memory loss walking down that foggy hill — at every single step, they feel the slope right under their feet and take a step proportional to it, then completely forget everything about the step they just took. If the ground is a narrow, steep-walled ravine, they'll bounce back and forth off the walls instead of walking smoothly down the middle.

What it does: at every step, plain stochastic gradient descent multiplies the current gradient by a learning rate and subtracts that from the current parameter value, with no memory of any earlier step at all.

Why it's needed: it's the simplest possible instantiation of gradient-based learning and remains the conceptual and computational baseline that every other optimizer on this list is compared against.

What fails without more than this: on loss surfaces shaped like narrow valleys — common in real neural network loss landscapes, where curvature differs sharply across directions — plain SGD oscillates back and forth across the steep direction while making frustratingly slow progress along the shallow direction toward the actual minimum.

🎯 Use this as the mental reference point before layering on momentum or adaptive learning rates.

3. Momentum: Remembering the Direction You Were Already Going

🧠 Analogy: now imagine a heavy bowling ball rolling down that same foggy hill instead of a person taking discrete steps. The ball has real momentum — a bump that would instantly redirect a light pebble barely nudges the ball off its established path, because most of its motion comes from the velocity it has already built up, not from the immediate slope alone.

What it does: momentum keeps a running "velocity" that blends the current gradient with a decayed version of the velocity from the previous step, and updates parameters using that blended velocity instead of the raw gradient alone. PyTorch's own SGD documentation writes the momentum update directly as v(t+1) = mu × v(t) + g(t+1), followed by p(t+1) = p(t) − lr × v(t+1), where p, g, v, and mu denote the parameter, gradient, velocity, and momentum factor respectively. (PyTorch SGD documentation)

Why it's needed: by accumulating velocity, momentum averages out the oscillating, back-and-forth components of the gradient across a narrow valley's steep walls, while consistently reinforcing the components pointing in the same direction step after step — which is exactly the shallow-but-persistent direction toward the true minimum.

How it works, step by step, for one parameter with momentum factor mu:

  1. Compute the current gradient g for this step.
  2. Update the velocity: multiply the previous velocity by mu, and add the current gradient.
  3. Update the parameter: subtract the learning rate times the new velocity from the current parameter value.
  4. Carry the updated velocity forward into the next step, so its influence decays geometrically rather than disappearing all at once.

What fails without it: without momentum, every step is decided from scratch using only the current gradient, so the optimizer has no way to "smooth over" a noisy or oscillating gradient signal — a problem that gets worse as loss surfaces become more elongated or as mini-batch gradient noise increases.

💡 Trade-off: PyTorch's SGD documentation explicitly notes that its implementation of momentum subtly differs from the formulation used in some other frameworks and in the classical Sutskever et al. formulation, so a momentum value tuned for one framework's implementation may not translate exactly to another's without adjustment.

🎯 Use this when a plain-SGD training run oscillates visibly on the loss curve or converges far more slowly than the loss surface's overall shape should require.

4. Adam: Adaptive Per-Parameter Learning Rates

🧠 Analogy: imagine that same rolling ball, except now it's actually a small robot with independently adjustable legs — one for every parameter — and each leg learns its own personal stride length based on how consistently and how strongly that specific leg has needed to move in the past. Legs that have been making big, decisive moves take smaller, more careful steps going forward; legs that have barely moved take bigger, bolder ones.

What it does: Adam, introduced by Diederik Kingma and Jimmy Ba in their 2014 paper "Adam: A Method for Stochastic Optimization," maintains two running averages for every parameter: a running average of the gradient itself (the first moment, playing a role similar to momentum) and a running average of the squared gradient (the second moment, tracking how large that parameter's gradients have typically been). Each parameter's effective learning rate is then scaled down by the square root of its own accumulated squared-gradient average. (Kingma & Ba, "Adam: A Method for Stochastic Optimization," 2014/ICLR 2015)

Why it's needed: different parameters in a large network can have wildly different gradient scales and sparsity patterns; a single global learning rate that works well for one parameter can be far too large or far too small for another. Adam's per-parameter adaptive scaling largely removes the need to hand-tune a single learning rate that works uniformly across every parameter in the network.

How it works, step by step, for one parameter:

  1. Compute the current gradient g.
  2. Update the first moment estimate: blend the previous first moment with the current gradient using decay rate beta1.
  3. Update the second moment estimate: blend the previous second moment with the current squared gradient using decay rate beta2.
  4. Apply a bias correction to both moment estimates, which compensates for both being initialized at zero and therefore biased toward zero in the earliest training steps.
  5. Update the parameter by subtracting the learning rate times the corrected first moment, divided by the square root of the corrected second moment plus a small stability constant epsilon.

What fails without understanding it: because Adam divides by the square root of an accumulated squared-gradient average, any additional term added directly into the gradient — including a naive L2 weight-decay penalty — gets divided by that same adaptive denominator, meaning its effective regularization strength ends up depending on each parameter's gradient history rather than being applied uniformly, as later work specifically identified.

🎯 Use this when training a network with highly heterogeneous gradient scales across parameters, such as models mixing embeddings with dense layers.

5. AdamW: Decoupling Weight Decay from the Adaptive Step

Two parallel flow diagrams starting from the same gradient. The top flow, labeled Adam with L2 regularization, adds the weight decay term into the gradient before it passes through the adaptive moment scaling step, so the decay gets divided by the adaptive denominator. The bottom flow, labeled AdamW, keeps the adaptive moment scaling step working on the raw gradient alone, and instead subtracts the weight decay term directly from the parameter after that step.

Figure 2. Original diagram: the only structural change AdamW makes is moving weight decay outside the adaptive step.

What it does: Ilya Loshchilov and Frank Hutter's paper "Decoupled Weight Decay Regularization" demonstrates that L2 regularization and weight decay, which are mathematically equivalent for plain SGD, are not equivalent for adaptive gradient algorithms like Adam, and proposes recovering true weight decay by applying it directly to the parameters, separately from the gradient-based optimization step. (Loshchilov & Hutter, "Decoupled Weight Decay Regularization," ICLR 2019) PyTorch's own AdamW documentation confirms this directly, describing the algorithm as one "where weight decay does not accumulate in the momentum nor variance." (PyTorch AdamW documentation)

Why it's needed: the paper's own abstract reports that this decoupling substantially improves Adam's generalization performance, allowing it to compete with SGD with momentum on image classification tasks where Adam had previously typically been outperformed by SGD with momentum. (Loshchilov & Hutter, 2019)

How it works, step by step, for one parameter:

  1. Compute the current gradient g using only the primary loss function — no weight-decay term mixed in.
  2. Run the standard Adam first- and second-moment update and bias correction exactly as described above, using this clean gradient.
  3. Compute the adaptive Adam update from those corrected moments as usual.
  4. Separately, subtract the weight decay coefficient times the current parameter value, scaled by the learning rate, directly from the parameter — a step that never passes through the adaptive denominator at all.

✅ Worked example: the abstract of the Decoupled Weight Decay Regularization paper reports that its proposed modification decouples the optimal choice of weight decay factor from the setting of the learning rate for both standard SGD and Adam — meaning that under the decoupled formulation, changing the learning rate no longer silently changes the effective regularization strength as well.

What fails without understanding it: using Adam with naive L2 regularization and expecting it to behave like traditional weight decay leads to regularization strength that varies unpredictably by parameter and by learning rate, which the Decoupled Weight Decay Regularization paper's abstract identifies as the exact inequivalence its method was designed to fix.

🎯 Use this when choosing between torch.optim.Adam and torch.optim.AdamW, or when a model regularized with Adam and L2 weight decay isn't generalizing as expected.

6. Implementation Walkthrough

The short, illustrative examples below show the concepts above as code. They are simplified teaching snippets, not production-ready modules.

# Illustrative example: momentum update, from scratch, for one parameter
velocity = 0.0

def momentum_step(param, grad, velocity, lr, mu):
    velocity = mu * velocity + grad
    param = param - lr * velocity
    return param, velocity

param, velocity = momentum_step(param, grad, velocity, lr=0.01, mu=0.9)
# Illustrative example: choosing an optimizer in PyTorch-style code
optimizer = create_optimizer(
    "AdamW",
    model.parameters(),
    lr=3e-4,
    betas=(0.9, 0.999),
    weight_decay=0.01,
)

for batch in data_loader:
    optimizer.zero_grad()
    loss = compute_loss(model, batch)
    loss.backward()
    optimizer.step()

Best practices: when switching from Adam to AdamW, re-tune the weight decay coefficient rather than reusing the same value, since the decoupled version applies it differently and the two are not interchangeable at the same numeric setting; and log each parameter group's effective learning rate and weight decay separately when using parameter groups that exclude biases or normalization parameters from decay.

🎯 Use this when wiring up a new training loop and deciding which optimizer and hyperparameters to start from.

7. Enterprise Rollout: Choosing, Testing, and Monitoring Optimizers

The optimizer and its hyperparameters are as much a part of a reproducible training pipeline as the model architecture itself, and should be governed with the same rigor.

Ownership and governance: document the exact optimizer, learning rate schedule, momentum or beta values, and weight decay setting used for every production training run, since these interact with each other in ways that make an isolated later change (like bumping the learning rate) potentially invalidate the tuning of another (like weight decay under AdamW).

CI gates: add a short regression test that trains a small model for a fixed number of steps on a fixed seed and dataset slice, and asserts the resulting loss falls within an expected range, so a refactor that accidentally changes optimizer state handling (like a momentum buffer not being reset correctly) is caught immediately rather than surfacing as unexplained training instability later.

Checkpoint versioning: save optimizer state (momentum buffers, or Adam's first and second moment estimates) alongside model weights in every checkpoint intended for resuming training, since restarting from weights alone without that state effectively restarts the optimizer's memory from zero and can visibly disrupt convergence right after a resume.

Dashboards and alerts: track gradient norm and loss curve smoothness over the course of training, and alert on sudden loss spikes or divergence, since these are frequently caused by a learning rate that's too aggressive for the chosen optimizer rather than a data or architecture problem.

Incident response: when a training run diverges or produces NaN losses, capture the exact optimizer type, learning rate, and step count at which it happened as part of the incident record, since an optimizer-related divergence is often reproducible from just those three facts without needing to replay the entire dataset.

🎯 Use this when a training pipeline built by one team is being handed off to another team that will retrain or fine-tune it regularly.

8. Common Mistakes

Using Adam with L2 weight decay and expecting it to behave like traditional weight decay. As the Decoupled Weight Decay Regularization paper's abstract demonstrates, this equivalence only holds for plain SGD, not for adaptive algorithms like Adam. The production impact is regularization that behaves inconsistently across parameters with different gradient histories, making weight decay tuning far less predictable than expected.

Reusing an Adam-tuned weight decay value directly with AdamW, or vice versa. Because the two apply weight decay through structurally different paths, the same numeric weight decay setting does not produce the same regularization strength in both. The production impact is a model that appears under-regularized or over-regularized purely because of a hyperparameter that was carried over without re-tuning.

Discarding optimizer state when resuming training from a checkpoint. Momentum buffers and Adam's moment estimates represent accumulated information about the training trajectory; restarting them from zero after a resume effectively restarts the optimizer's "memory," which can produce a visible disruption in the loss curve right at the resume point.

Assuming a higher momentum or a higher Adam beta value is always better. Both control how much of the past is remembered; set too high, they can make the optimizer sluggish to respond to genuine, recent changes in the loss landscape, effectively trading responsiveness for smoothness rather than improving both.

Copying a framework's default optimizer implementation detail across frameworks without checking. PyTorch's own SGD documentation explicitly notes its momentum implementation subtly differs from other formulations and other frameworks; assuming identical numerical behavior when porting a training recipe between frameworks can silently change results in ways that are hard to attribute to the actual cause.

❓ FAQ

Is Adam always a better choice than SGD with momentum?

Not universally. The Decoupled Weight Decay Regularization paper's own abstract notes that Adam was previously typically outperformed by SGD with momentum on certain image classification tasks, and that decoupled weight decay was specifically what let Adam compete on that front — so the "better" choice depends on the task, architecture, and how carefully each optimizer's hyperparameters are tuned.

What's the actual difference between Adam and AdamW in one sentence?

AdamW applies weight decay directly to the parameters after the adaptive update step, while Adam with L2 regularization mixes the decay term into the gradient before that same adaptive step, causing it to be scaled unevenly.

Does momentum ever hurt training?

It can if set too high for the loss landscape's curvature, since a large accumulated velocity can carry the optimizer past a minimum it was approaching, causing oscillation around it rather than a clean convergence — this is why momentum is a tunable hyperparameter rather than a fixed constant.

Why does Adam need two running averages instead of just one?

The first moment (average of the gradient) provides the momentum-like directional smoothing, while the second moment (average of the squared gradient) is what lets Adam compute a separate adaptive learning rate for every parameter based on how large that parameter's gradients have typically been.

Can I just always leave optimizer hyperparameters at their framework defaults?

Defaults are reasonable, well-tested starting points, but the Decoupled Weight Decay Regularization paper's central finding is specifically that the interaction between learning rate and weight decay changes across optimizer formulations — so at minimum, weight decay and learning rate should be reconsidered together, not tuned as if they were independent of the optimizer choice.

🔗 References & Further Reading

📝 Summary

  • Every optimizer here consumes the same backpropagated gradients; they differ only in how those gradients become a weight update.
  • Plain SGD has no memory between steps and can oscillate badly on narrow, ill-conditioned loss surfaces.
  • Momentum accumulates a decaying velocity from past gradients, smoothing out oscillation while reinforcing consistent directions.
  • Adam adds a second, independent running average per parameter, giving every parameter its own adaptive effective learning rate.
  • AdamW decouples weight decay from the adaptive update entirely, fixing a documented inequivalence between L2 regularization and true weight decay under Adam.
  • Treat the full optimizer configuration — algorithm, learning rate, momentum or betas, and weight decay — as one versioned, tested, monitored unit, not independent settings.

From a forgetful hiker to a rolling ball to a robot with independently smart legs — that's the whole optimizer lineage. Happy training! 🚀

Comments