COURSE 05L1100% FREE
Verified 2026-08-10

Optimizers: SGD and Adam

An **optimizer** decides how to nudge weights using gradients. **SGD** steps opposite the gradient (often with minibatches and momentum). **Adam** adapts step sizes using running averages of gradient statistics.

What it is

An optimizer decides how to nudge weights using gradients. SGD steps opposite the gradient (often with minibatches and momentum). Adam adapts step sizes using running averages of gradient statistics.

Why it matters

Learning rate and optimizer choice often matter as much as architecture for “does training work?”

How it works (plain)

Gradient says “uphill for loss.” Optimizers walk downhill with different gaits: steady small steps, momentum that builds speed, or adaptive steps per parameter (Adam family).

Everyday example

Hiking downhill: careful steps (small LR) vs long strides that overshoot; Adam is like adjusting stride per terrain type.

Try it

Train a mental model: if loss explodes, what do you try first (usually lower LR / check bugs)?

Myths

⚠️ Myth: Adam always beats SGD.
✓ Reality: Task-dependent; many vision recipes still favor SGD+momentum carefully tuned.
⚠️ Myth: Default hyperparameters are universal.
✓ Reality: Defaults are starting points—validate.

Sources