COURSE 04L2100% FREE
Verified 2026-08-10

Derivatives and chain-rule intuition

A **derivative** measures how fast an output changes when an input nudges a little—slope. The **chain rule** says: when functions stack, slopes multiply along the path.

What it is

A derivative measures how fast an output changes when an input nudges a little—slope. The chain rule says: when functions stack, slopes multiply along the path.

Why it matters

Training neural nets is mostly “nudge weights to reduce error.” Backprop is organized chain-rule bookkeeping (Course 05).

How it works (plain)

If loss depends on layer 3, which depends on layer 2, which depends on a weight, you multiply the local slopes to learn how that weight affects the loss. Steep slope → bigger suggested nudge (before optimizers soften it).

Everyday example

Volume dial → amp → speaker loudness. A small dial turn can cause a big loudness change if each stage amplifies—slopes compose.

Try it

Sketch a two-step machine: input → double → add 1. If input rises by 0.1, how much does output rise?

Myths

⚠️ Myth: You must hand-derive every network to use deep learning.
✓ Reality: Autodiff libraries do the bookkeeping; intuition still prevents misuse.
⚠️ Myth: Zero gradient always means “done.”
✓ Reality: It can mean stuck, saturated activations, or a bug.

Sources