Derivatives and chain-rule intuition
A **derivative** measures how fast an output changes when an input nudges a little—slope. The **chain rule** says: when functions stack, slopes multiply along the path.
What it is
A derivative measures how fast an output changes when an input nudges a little—slope. The chain rule says: when functions stack, slopes multiply along the path.
Why it matters
Training neural nets is mostly “nudge weights to reduce error.” Backprop is organized chain-rule bookkeeping (Course 05).
How it works (plain)
If loss depends on layer 3, which depends on layer 2, which depends on a weight, you multiply the local slopes to learn how that weight affects the loss. Steep slope → bigger suggested nudge (before optimizers soften it).
Everyday example
Volume dial → amp → speaker loudness. A small dial turn can cause a big loudness change if each stage amplifies—slopes compose.
Try it
Sketch a two-step machine: input → double → add 1. If input rises by 0.1, how much does output rise?
Myths
- ⚠️ Myth: You must hand-derive every network to use deep learning.
- ✓ Reality: Autodiff libraries do the bookkeeping; intuition still prevents misuse.
- ⚠️ Myth: Zero gradient always means “done.”
- ✓ Reality: It can mean stuck, saturated activations, or a bug.
Sources
- 3Blue1Brown calculus / neural net intros: https://www.3blue1brown.com/ ↗
- Course 05 backpropagation
- fast.ai practical stance: https://www.fast.ai/ ↗
