COURSE 18L1100% FREE
Verified 2026-08-14

Task success vs proxy metrics

Separating **true task success** (did the user finish the job safely?) from **proxy metrics** (BLEU, thumbs, latency, token count, MMLU) that only partially correlate. OpenAI’s GDPval work (2025) pushes toward economically valuable knowl...

What it is

Separating true task success (did the user finish the job safely?) from proxy metrics (BLEU, thumbs, latency, token count, MMLU) that only partially correlate. OpenAI’s GDPval work (2025) pushes toward economically valuable knowledge-work tasks—closer to real jobs than pure quiz proxies.

MEDIUM PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Proxy vs task success: BLEU/accuracy gauges vs user task completed checkbox; Goodhart warning.

Educational Focus: Metric literacy punchline.

Why it matters

Teams ship to the proxy and miss the mission—classic Goodhart / reward-hacking cousin (Course 14). Anthropic warns that a single accuracy number can hide useless “unbiased” zeros when models refuse everything.

How it works (plain)

Define the job → pick a gold success definition → use proxies for speed → periodically validate proxies against gold → never optimize a proxy alone for high stakes.

Everyday example

Bathroom scale weight is a proxy for health—useful weekly, misleading as the only medical signal.

Try it

For one AI feature, write the gold success definition and two proxies—with one known failure mode of each proxy.

Myths

⚠️ Myth: If all proxies are green, users are happy.
✓ Reality: Proxies lag reality—sample real tasks.
⚠️ Myth: Human preference always equals task success.
✓ Reality: Preference can reward style over correctness—calibrate.

Sources