Weight initialization
How you **set starting weights** before training. Bad init can stall learning; good init keeps signals and gradients in a healthy range.
What it is
How you set starting weights before training. Bad init can stall learning; good init keeps signals and gradients in a healthy range.
Why it matters
Explains mysterious “my net won’t train” failures alongside LR and architecture choices.
How it works (plain)
Too large → explosions; too small → silence. Modern schemes scale random weights by layer size (Glorot/He families). Frameworks often default sensibly—know what they assume (activation type).
Everyday example
Tuning guitar strings before a song—not random yanks.
Try it
When a lab fails to learn, check init + activation pairing before rewriting the model.
Myths
- ⚠️ Myth: All zeros is a fine start.
- ✓ Reality: Symmetry breaking fails; units stay twins.
Sources
- Course 05 vanishing gradients; activations
- Glorot & Bengio / He et al. init papers (cite when teaching)
