COURSE 18L1100% FREE
Verified 2026-08-14

LLM-as-judge caveats

Using a model to **score** other model outputs—useful, biased, and gameable if you aren’t careful. Anthropic groups this with the broader challenge of evaluating generative systems: open-ended behavior is hard to reduce to a single number.

What it is

Using a model to score other model outputs—useful, biased, and gameable if you aren’t careful. Anthropic groups this with the broader challenge of evaluating generative systems: open-ended behavior is hard to reduce to a single number.

HIGH PRIORITYINFOGRAPHIC
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

LLM-as-judge caveats poster: position bias, self-preference, verbosity bias, rubric sensitivity, need human calibration. Good practice: blinded pairs, multiple judges, gold spots.

Educational Focus: Prevents naive judge pipelines.

Why it matters

Teams scale eval with judges. Without calibration, you automate preference for verbosity, shared quirks, or “sounds confident.” Human A/B preference tests (Anthropic) remain expensive but closer to real dialogue use.

How it works (plain)

Define rubric → blind pairs when possible → calibrate on human labels → watch position bias and self-preference → keep exact checkers for facts/formats → sample humans on high-severity cases.

Everyday example

A teacher using a rubric assistant still spot-checks grades.

Try it

Have a judge score two answers; swap order; see if rankings flip.

Myths

⚠️ Myth: Strong judges replace humans forever.
✓ Reality: Humans remain the anchor for high stakes.
⚠️ Myth: Same-family judge is neutral.
✓ Reality: Self-preference effects appear in research—test for them.
⚠️ Myth: Model-written evals need no human check.
✓ Reality: Anthropic’s model-written evaluations still rely on humans to verify accuracy and labels.

Sources