COURSE 05L1100% FREE
Verified 2026-08-10

Vanishing and exploding gradients

When training deep or recurrent nets, gradients can **shrink toward zero** (vanishing) or **blow up** (exploding) as they pass through many layers/time steps—blocking useful learning.

What it is

When training deep or recurrent nets, gradients can shrink toward zero (vanishing) or blow up (exploding) as they pass through many layers/time steps—blocking useful learning.

Why it matters

Explains why activations, initialization, residuals, and clipping exist—and why early deep nets were hard to train.

How it works (plain)

Backprop multiplies many local slopes. Numbers <1 multiply toward 0; numbers >1 can explode. Saturated activations and deep stacks amplify the problem.

Everyday example

A whisper passed through twenty people—or a rumor shouted each time—signal dies or becomes noise.

Try it

Multiply 0.9 by itself ten times vs 1.1 ten times on a calculator—feel vanish vs explode.

Myths

⚠️ Myth: Only RNNs have this problem.
✓ Reality: Deep feedforward nets suffered too; residuals/normalization helped.
⚠️ Myth: Exploding loss always means bad data.
✓ Reality: Often LR, init, or bugs—inspect gradients.

Sources