Optimizers: SGD and Adam
An **optimizer** decides how to nudge weights using gradients. **SGD** steps opposite the gradient (often with minibatches and momentum). **Adam** adapts step sizes using running averages of gradient statistics.
What it is
An optimizer decides how to nudge weights using gradients. SGD steps opposite the gradient (often with minibatches and momentum). Adam adapts step sizes using running averages of gradient statistics.
Why it matters
Learning rate and optimizer choice often matter as much as architecture for “does training work?”
How it works (plain)
Gradient says “uphill for loss.” Optimizers walk downhill with different gaits: steady small steps, momentum that builds speed, or adaptive steps per parameter (Adam family).
Everyday example
Hiking downhill: careful steps (small LR) vs long strides that overshoot; Adam is like adjusting stride per terrain type.
Try it
Train a mental model: if loss explodes, what do you try first (usually lower LR / check bugs)?
Myths
- ⚠️ Myth: Adam always beats SGD.
- ✓ Reality: Task-dependent; many vision recipes still favor SGD+momentum carefully tuned.
- ⚠️ Myth: Default hyperparameters are universal.
- ✓ Reality: Defaults are starting points—validate.
Sources
- Course 05 backpropagation; Course 04 calculus intuition
- Kingma & Ba Adam paper: https://arxiv.org/abs/1412.6980 ↗
- fast.ai: https://www.fast.ai/ ↗
