Batch norm vs layer norm (deeper)
A deeper comparison of **batch normalization** vs **layer normalization** (and friends): what statistics they use and when each shows up.
What it is
A deeper comparison of batch normalization vs layer normalization (and friends): what statistics they use and when each shows up.
Why it matters
Vision CNNs historically loved batch norm; transformers often use layer/RMS norm. Mixing them blindly breaks recipes.
How it works (plain)
Batch norm: normalize across the batch for each channel/feature (train vs eval behavior differs). Layer norm: normalize across features for each example—friendlier for variable NLP batches and decoders.
Everyday example
Grading on a curve within a class (batch-ish) vs standardizing each student’s own answer vector (layer-ish)—imperfect metaphor, useful for memory.
Try it
When reading a paper diagram, label which norm sits where.
Myths
- ⚠️ Myth: Norm layers are optional cosmetics.
- ✓ Reality: They often decide whether deep stacks train at all.
Sources
- Course 05 residuals-and-normalization; Course 07 transformers
- Ioffe & Szegedy BatchNorm; Ba et al. LayerNorm (cite when teaching)
