Entropy and cross-entropy
**Entropy** measures uncertainty in a distribution—how surprising outcomes are on average. **Cross-entropy** measures how well one distribution (a model’s predictions) matches another (the true labels)—and is a common classification loss.
What it is
Entropy measures uncertainty in a distribution—how surprising outcomes are on average. Cross-entropy measures how well one distribution (a model’s predictions) matches another (the true labels)—and is a common classification loss.
Why it matters
When people say a model “minimizes cross-entropy,” they mean it nudges predicted probabilities toward the correct classes. This links math to training loops.
How it works (plain)
A fair coin has higher entropy than a coin that almost always lands heads. If your model is confident and wrong, cross-entropy punishes hard; if it hedges appropriately, loss is gentler.
Everyday example
A weather app that always says “50% rain” vs one that matches real frequencies—calibration and surprise differ.
Try it
List three everyday uncertain events and rank them from low to high surprise-when-they-happen.
Myths
- ⚠️ Myth: Lowest training loss always means best real-world model.
- ✓ Reality: Overfitting and metric mismatch still happen.
- ⚠️ Myth: Entropy means “information” in the everyday news sense.
- ✓ Reality: It is a precise uncertainty measure in bits/nats.
Sources
- Course 03 features-loss-and-optimization; Course 07 sampling
- deep learning book info theory notes: https://www.deeplearningbook.org/ ↗
- 3Blue1Brown / visual explainers (optional)
