COURSE 02L0100% FREE
Verified 2026-08-10

Data and labels

Most modern AI systems learn from **examples**. **Data** is those examples. **Labels** are the “answer keys” attached to some examples (spam/not spam, price, diagnosis code, next word). Messy data makes messy models. Biased labels replay...

What it is

Most modern AI systems learn from examples. Data is those examples. Labels are the “answer keys” attached to some examples (spam/not spam, price, diagnosis code, next word).

Messy data makes messy models. Biased labels replay human bias (Course 19).

Why it matters

People argue about models. Pros often argue about datasets. If you only change the algorithm and ignore the data, you are guessing.

How it works (plain)

  • Collect inputs (emails, photos, sensor rows, documents)
  • Optionally attach labels
  • Clean, split, and document what the data covers—and what it misses
  • Train, then watch for drift when the real world changes

Everyday example

A résumé screener trained mostly on past hires from one demographic may underrate other qualified people—not because “math is evil,” but because history was uneven.

Try it

Pick a prediction you care about. Write: (1) what inputs you would collect, (2) who would label them, (3) how disagreements get resolved.

Myths

⚠️ Myth: More data always fixes the model.
✓ Reality: More *of the same blind spot* can lock the blind spot in.
⚠️ Myth: Labels are objective truth.
✓ Reality: Many labels are judgments; guidelines matter.

Sources