Task success vs proxy metrics
Separating **true task success** (did the user finish the job safely?) from **proxy metrics** (BLEU, thumbs, latency, token count, MMLU) that only partially correlate. OpenAI’s GDPval work (2025) pushes toward economically valuable knowl...
What it is
Separating true task success (did the user finish the job safely?) from proxy metrics (BLEU, thumbs, latency, token count, MMLU) that only partially correlate. OpenAI’s GDPval work (2025) pushes toward economically valuable knowledge-work tasks—closer to real jobs than pure quiz proxies.
Visual Spec & Architecture Diagram
Proxy vs task success: BLEU/accuracy gauges vs user task completed checkbox; Goodhart warning.
Why it matters
Teams ship to the proxy and miss the mission—classic Goodhart / reward-hacking cousin (Course 14). Anthropic warns that a single accuracy number can hide useless “unbiased” zeros when models refuse everything.
How it works (plain)
Define the job → pick a gold success definition → use proxies for speed → periodically validate proxies against gold → never optimize a proxy alone for high stakes.
Everyday example
Bathroom scale weight is a proxy for health—useful weekly, misleading as the only medical signal.
Try it
For one AI feature, write the gold success definition and two proxies—with one known failure mode of each proxy.
Myths
- ⚠️ Myth: If all proxies are green, users are happy.
- ✓ Reality: Proxies lag reality—sample real tasks.
- ⚠️ Myth: Human preference always equals task success.
- ✓ Reality: Preference can reward style over correctness—calibrate.
Sources
- OpenAI Evals / GDPval: https://evals.openai.com/ ↗
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
- Course 14 reward hacking; Course 01 measuring progress
