Vanishing and exploding gradients
When training deep or recurrent nets, gradients can **shrink toward zero** (vanishing) or **blow up** (exploding) as they pass through many layers/time steps—blocking useful learning.
What it is
When training deep or recurrent nets, gradients can shrink toward zero (vanishing) or blow up (exploding) as they pass through many layers/time steps—blocking useful learning.
Why it matters
Explains why activations, initialization, residuals, and clipping exist—and why early deep nets were hard to train.
How it works (plain)
Backprop multiplies many local slopes. Numbers <1 multiply toward 0; numbers >1 can explode. Saturated activations and deep stacks amplify the problem.
Everyday example
A whisper passed through twenty people—or a rumor shouted each time—signal dies or becomes noise.
Try it
Multiply 0.9 by itself ten times vs 1.1 ten times on a calculator—feel vanish vs explode.
Myths
- ⚠️ Myth: Only RNNs have this problem.
- ✓ Reality: Deep feedforward nets suffered too; residuals/normalization helped.
- ⚠️ Myth: Exploding loss always means bad data.
- ✓ Reality: Often LR, init, or bugs—inspect gradients.
Sources
- Course 05 residuals; RNNs; backpropagation
- Glorot/He initialization literature (cite when teaching)
- deep learning book: https://www.deeplearningbook.org/ ↗
