Datasets and splits
A **dataset** is the collection of examples a model learns from. A **split** divides that collection into parts for training, checking, and final testing so you do not grade the model on the same homework it memorized.
What it is
A dataset is the collection of examples a model learns from. A split divides that collection into parts for training, checking, and final testing so you do not grade the model on the same homework it memorized.
Why it matters
Bad splits create fake success. If test examples leaked into training, scores look great and real-world use fails.
How it works (plain)
Common split:
- Train: learn patterns
- Validation: tune choices (model size, learning rate)
- Test: final report card—touch once
Keep people, time periods, or sites from leaking across splits when that matches reality (e.g., don’t put the same patient’s visits in train and test for a medical demo).
Everyday example
Studying with the answer key in front of you, then taking a “test” that is the same worksheet—looks like mastery, isn’t.
Try it
For a CSV you care about, write how you would split it and what leakage you fear.
Myths
- ⚠️ Myth: Random 80/20 always works.
- ✓ Reality: Time series and grouped data need careful splits.
- ⚠️ Myth: Bigger datasets automatically fix bias.
- ✓ Reality: Bigger can amplify the same skew.
Sources
- fast.ai: https://www.fast.ai/ ↗
- Course 02 data-and-labels; Course 03 generalization
- Google ML Crash Course (data): https://developers.google.com/machine-learning/crash-course ↗
