Learning rate schedules
Plans for changing the **learning rate** during training: warmups, decays, cosine schedules, plateaus.
What it is
Plans for changing the learning rate during training: warmups, decays, cosine schedules, plateaus.
Why it matters
A fixed LR is often suboptimal. Schedules are standard knobs in serious training recipes.
How it works (plain)
Start careful (warmup) → train → reduce LR to settle into a sharper minimum region (intuitive story). Watch val curves when you drop LR.
Everyday example
Sprinting then jogging the last mile—pace changes with the goal.
Try it
On a tiny lab, compare constant LR vs a single drop halfway—note val loss.
Myths
- ⚠️ Myth: Lower LR is always safer and better.
- ✓ Reality: Too low underfits or crawls forever.
Sources
- Course 05 optimizers; debugging training curves
- fast.ai: https://www.fast.ai/ ↗
