Metrics: precision, recall, and ROC
Beyond accuracy: **precision** (of predicted positives, how many were right), **recall** (of real positives, how many you found), and curves like **ROC** that show tradeoffs as you change thresholds.
What it is
Beyond accuracy: precision (of predicted positives, how many were right), recall (of real positives, how many you found), and curves like ROC that show tradeoffs as you change thresholds.
Why it matters
Imbalanced problems make accuracy misleading. Choosing metrics is choosing whose errors matter.
How it works (plain)
Spam filter: high precision → fewer false spam tags; high recall → fewer missed spam. You rarely max both without cost. Thresholds move the tradeoff.
Everyday example
Airport security vs email filters optimize different error costs.
Try it
For one problem, say which is worse: false positive or false negative—and pick a metric that reflects that.
Myths
- ⚠️ Myth: AUC solves metric choice forever.
- ✓ Reality: Useful summary; still check operating points and calibration.
- ⚠️ Myth: Accuracy is fine if it’s “high.”
- ✓ Reality: 99% accuracy on 1% rare events can be trivial.
Sources
- Course 03 supervised learning; Course 04 statistics
- Google ML Crash Course classification: https://developers.google.com/machine-learning/crash-course ↗
- scikit-learn metrics docs: https://scikit-learn.org/ ↗
