Contamination awareness
When eval items (or very close paraphrases) appear in training data—so scores look better than true generalization. Anthropic’s evaluation challenges essay compares this to students seeing the test questions beforehand.
What it is
When eval items (or very close paraphrases) appear in training data—so scores look better than true generalization. Anthropic’s evaluation challenges essay compares this to students seeing the test questions beforehand.
Visual Spec & Architecture Diagram
Contamination diagram: benchmark items leaking into pretraining data lake; test score inflated gauge; mitigations: canaries, held-out, dynamic exams. Title: 'Train/test contamination'.
Why it matters
Leaderboards can lie. Buyers and scientists need contamination literacy before trusting sudden jumps on public quizzes.
How it works (plain)
Prefer held-out private tasks; canary strings; overlap searches when possible; distrust sudden jumps on popular public quizzes without methods; refresh private suites periodically.
Everyday example
If the answer key leaked into the study guide, the exam score stops meaning “learned the subject.”
Try it
Create a private 20-item eval with unique canaries you never publish.
Myths
- ⚠️ Myth: Big public benchmarks are automatically clean.
- ✓ Reality: Leak risk rises with popularity and time (Anthropic on MMLU).
- ⚠️ Myth: Private evals never leak.
- ✓ Reality: Screenshots, support tickets, and vendor demos can still expose items—control distribution.
Sources
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
- HELM Capabilities (transparent, reproducible scenarios): https://crfm.stanford.edu/helm/capabilities/v1.14.0/ ↗
- Course 01 measuring-progress; Course 23 papers literacy
