Data and labels
Most modern AI systems learn from **examples**. **Data** is those examples. **Labels** are the “answer keys” attached to some examples (spam/not spam, price, diagnosis code, next word). Messy data makes messy models. Biased labels replay...
What it is
Most modern AI systems learn from examples. Data is those examples. Labels are the “answer keys” attached to some examples (spam/not spam, price, diagnosis code, next word).
Messy data makes messy models. Biased labels replay human bias (Course 19).
Why it matters
People argue about models. Pros often argue about datasets. If you only change the algorithm and ignore the data, you are guessing.
How it works (plain)
- Collect inputs (emails, photos, sensor rows, documents)
- Optionally attach labels
- Clean, split, and document what the data covers—and what it misses
- Train, then watch for drift when the real world changes
Everyday example
A résumé screener trained mostly on past hires from one demographic may underrate other qualified people—not because “math is evil,” but because history was uneven.
Try it
Pick a prediction you care about. Write: (1) what inputs you would collect, (2) who would label them, (3) how disagreements get resolved.
Myths
- ⚠️ Myth: More data always fixes the model.
- ✓ Reality: More *of the same blind spot* can lock the blind spot in.
- ⚠️ Myth: Labels are objective truth.
- ✓ Reality: Many labels are judgments; guidelines matter.
Sources
- Google ML Crash Course (data/labels framing): https://developers.google.com/machine-learning/crash-course ↗
- ANN live: https://www.ainerdnetwork.com/learn/data-and-labels ↗
- NIST AI RMF (map/measure data risks): https://www.nist.gov/itl/ai-risk-management-framework ↗
