LLM-as-judge caveats
Using a model to **score** other model outputs—useful, biased, and gameable if you aren’t careful. Anthropic groups this with the broader challenge of evaluating generative systems: open-ended behavior is hard to reduce to a single number.
What it is
Using a model to score other model outputs—useful, biased, and gameable if you aren’t careful. Anthropic groups this with the broader challenge of evaluating generative systems: open-ended behavior is hard to reduce to a single number.
Visual Spec & Architecture Diagram
LLM-as-judge caveats poster: position bias, self-preference, verbosity bias, rubric sensitivity, need human calibration. Good practice: blinded pairs, multiple judges, gold spots.
Why it matters
Teams scale eval with judges. Without calibration, you automate preference for verbosity, shared quirks, or “sounds confident.” Human A/B preference tests (Anthropic) remain expensive but closer to real dialogue use.
How it works (plain)
Define rubric → blind pairs when possible → calibrate on human labels → watch position bias and self-preference → keep exact checkers for facts/formats → sample humans on high-severity cases.
Everyday example
A teacher using a rubric assistant still spot-checks grades.
Try it
Have a judge score two answers; swap order; see if rankings flip.
Myths
- ⚠️ Myth: Strong judges replace humans forever.
- ✓ Reality: Humans remain the anchor for high stakes.
- ⚠️ Myth: Same-family judge is neutral.
- ✓ Reality: Self-preference effects appear in research—test for them.
- ⚠️ Myth: Model-written evals need no human check.
- ✓ Reality: Anthropic’s model-written evaluations still rely on humans to verify accuracy and labels.
Sources
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
- Anthropic — Model-Written Evaluations: https://www.anthropic.com/research/discovering-language-model-behaviors-with-model-written-evaluations ↗
- Course 18 overview; Course 10 agent eval
