Activation functions
**Activations** are simple nonlinear functions applied between layers so networks can learn curved decision boundaries—not only straight lines.
What it is
Activations are simple nonlinear functions applied between layers so networks can learn curved decision boundaries—not only straight lines.
Why it matters
Without nonlinearities, deep stacks collapse to one linear map. Choice of activation affects training speed and gradient health.
How it works (plain)
ReLU family: pass positives, zero/suppress negatives (with variants). Sigmoid/tanh: squashing curves (historically common; can saturate). Modern nets often prefer ReLU-like or GELU in transformers.
Everyday example
A light dimmer that only responds above a click threshold—small signals ignored, larger ones pass.
Try it
Sketch y = max(0, x). Mark where gradients are zero vs one.
Myths
- ⚠️ Myth: Any nonlinearity is fine forever.
- ✓ Reality: Saturation and dying ReLUs are real training issues.
- ⚠️ Myth: Activations are the main source of “intelligence.”
- ✓ Reality: They’re necessary plumbing; data and objectives dominate.
Sources
- Course 05 neural-networks; vanishing gradients
- deep learning book: https://www.deeplearningbook.org/ ↗
