Leakage case studies
**Leakage** is when information from outside the training-time-allowed view sneaks into features or splits—so scores look great and production fails.
What it is
Leakage is when information from outside the training-time-allowed view sneaks into features or splits—so scores look great and production fails.
Why it matters
Leakage is one of the most common silent ML failures in industry tutorials and dashboards.
How it works (plain)
Classic patterns: target encoded with full-data stats; future timestamps in features; same user/patient in train and test; preprocessing fit on all rows; duplicates across splits.
Everyday example
Predicting “will buy?” with a feature that is only filled after purchase.
Try it
For a dataset you know, invent one leaky feature and one honest substitute.
Myths
- ⚠️ Myth: Random splits prevent all leakage.
- ✓ Reality: Group/time leakage survives randomness.
- ⚠️ Myth: If the pipeline is automated, leakage can’t happen.
- ✓ Reality: Automation can scale leakage.
Sources
- Course 02 datasets-and-splits; Course 03 generalization
- Google ML Crash Course data leakage notes: https://developers.google.com/machine-learning/crash-course ↗
