COURSE 18L1100% FREE
Verified 2026-08-14

Contamination awareness

When eval items (or very close paraphrases) appear in training data—so scores look better than true generalization. Anthropic’s evaluation challenges essay compares this to students seeing the test questions beforehand.

What it is

When eval items (or very close paraphrases) appear in training data—so scores look better than true generalization. Anthropic’s evaluation challenges essay compares this to students seeing the test questions beforehand.

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Contamination diagram: benchmark items leaking into pretraining data lake; test score inflated gauge; mitigations: canaries, held-out, dynamic exams. Title: 'Train/test contamination'.

Educational Focus: Explains why leaderboards lie.

Why it matters

Leaderboards can lie. Buyers and scientists need contamination literacy before trusting sudden jumps on public quizzes.

How it works (plain)

Prefer held-out private tasks; canary strings; overlap searches when possible; distrust sudden jumps on popular public quizzes without methods; refresh private suites periodically.

Everyday example

If the answer key leaked into the study guide, the exam score stops meaning “learned the subject.”

Try it

Create a private 20-item eval with unique canaries you never publish.

Myths

⚠️ Myth: Big public benchmarks are automatically clean.
✓ Reality: Leak risk rises with popularity and time (Anthropic on MMLU).
⚠️ Myth: Private evals never leak.
✓ Reality: Screenshots, support tickets, and vendor demos can still expose items—control distribution.

Sources