Debugging training curves
Reading **train vs validation loss/accuracy curves** to diagnose underfit, overfit, instability, and data bugs.
What it is
Reading train vs validation loss/accuracy curves to diagnose underfit, overfit, instability, and data bugs.
Why it matters
Curves are the X-ray of training. Most “mystery failures” show up here first.
How it works (plain)
Both high → underfit/capacity/LR/data. Train ↓ val ↑ → overfit. Wild spikes → LR/bugs/bad batches. Train not moving → dead net/wrong loss/labels.
Everyday example
A fitness tracker: if resting heart rate trends contradict workouts, investigate sensors before diets.
Try it
Sketch four curve patterns and label the first fix you’d try for each.
Myths
- ⚠️ Myth: Smooth curves mean a correct pipeline.
- ✓ Reality: Smooth wrong labels still look smooth.
Sources
- Course 03 overfitting; Course 05 regularization
- fast.ai: https://www.fast.ai/ ↗
