COURSE 05L1100% FREE
Verified 2026-08-10

Weight initialization

How you **set starting weights** before training. Bad init can stall learning; good init keeps signals and gradients in a healthy range.

What it is

How you set starting weights before training. Bad init can stall learning; good init keeps signals and gradients in a healthy range.

Why it matters

Explains mysterious “my net won’t train” failures alongside LR and architecture choices.

How it works (plain)

Too large → explosions; too small → silence. Modern schemes scale random weights by layer size (Glorot/He families). Frameworks often default sensibly—know what they assume (activation type).

Everyday example

Tuning guitar strings before a song—not random yanks.

Try it

When a lab fails to learn, check init + activation pairing before rewriting the model.

Myths

⚠️ Myth: All zeros is a fine start.
✓ Reality: Symmetry breaking fails; units stay twins.

Sources

  • Course 05 vanishing gradients; activations
  • Glorot & Bengio / He et al. init papers (cite when teaching)