COURSE 05L1100% FREE
Verified 2026-08-10

Residuals and normalization

**Residual connections** let a layer learn a change to add on top of its input (“skip connections”). **Normalization layers** re-center/re-scale activations so training stays stable.

What it is

Residual connections let a layer learn a change to add on top of its input (“skip connections”). Normalization layers re-center/re-scale activations so training stays stable.

Why it matters

These ideas helped make very deep networks trainable—and show up inside transformers too (residual streams around attention/MLP blocks).

How it works (plain)

Instead of forcing each layer to rebuild the whole representation, residuals say: “pass the old signal forward, and learn the delta.” Normalization keeps numbers from drifting into awkward ranges.

Everyday example

Editing a document with “track changes” (deltas) instead of retyping the whole file each time.

Try it

When you see a model diagram with arrows that skip boxes, label them “residual” and ask what still flows if the box learns ~0.

Myths

⚠️ Myth: Residuals mean the model isn’t deep anymore.
✓ Reality: Depth remains; gradients and information have a highway.
⚠️ Myth: All normalization layers behave the same.
✓ Reality: Batch vs layer vs RMSNorm differ—especially with small batches or transformers.

Sources