COURSE 02L0100% FREE
Verified 2026-08-10

Data quality and bias in labels

**Labels** are the answers attached to examples (spam/not spam, cat/dog). **Quality** means those answers are consistent, timely, and fair enough for the job. Bias in labels becomes bias in models.

What it is

Labels are the answers attached to examples (spam/not spam, cat/dog). Quality means those answers are consistent, timely, and fair enough for the job. Bias in labels becomes bias in models.

Why it matters

Models amplify what they are taught. If annotators disagree, or one group is systematically mislabeled, the system inherits that.

How it works (plain)

Watch for:

  • Ambiguous guidelines
  • Exhausted labelers rushing
  • Majority-vote hiding disagreement
  • Historical labels that encode past discrimination

Fix with clearer rules, second opinions, and measuring disagreement—not only more data.

Everyday example

Two moderators mark the same borderline comment differently every day—the “truth” is noisy.

Try it

Write a 5-bullet labeling guide for one task you care about. Where would two people disagree?

Myths

⚠️ Myth: Crowdsourcing is automatically unbiased.
✓ Reality: Who labels, and under what pay/time pressure, shapes results.
⚠️ Myth: Perfect labels exist for every social task.
✓ Reality: Some tasks are contested; document uncertainty.

Sources