Residuals and normalization
**Residual connections** let a layer learn a change to add on top of its input (“skip connections”). **Normalization layers** re-center/re-scale activations so training stays stable.
What it is
Residual connections let a layer learn a change to add on top of its input (“skip connections”). Normalization layers re-center/re-scale activations so training stays stable.
Why it matters
These ideas helped make very deep networks trainable—and show up inside transformers too (residual streams around attention/MLP blocks).
How it works (plain)
Instead of forcing each layer to rebuild the whole representation, residuals say: “pass the old signal forward, and learn the delta.” Normalization keeps numbers from drifting into awkward ranges.
Everyday example
Editing a document with “track changes” (deltas) instead of retyping the whole file each time.
Try it
When you see a model diagram with arrows that skip boxes, label them “residual” and ask what still flows if the box learns ~0.
Myths
- ⚠️ Myth: Residuals mean the model isn’t deep anymore.
- ✓ Reality: Depth remains; gradients and information have a highway.
- ⚠️ Myth: All normalization layers behave the same.
- ✓ Reality: Batch vs layer vs RMSNorm differ—especially with small batches or transformers.
Sources
- He et al., ResNet (2016) — https://arxiv.org/abs/1512.03385 ↗
- Course 05 neural-networks; Course 07 transformers
- deep learning book: https://www.deeplearningbook.org/ ↗
