COURSE 18L1100% FREE
Verified 2026-08-14

Model-written evals and HarmBench

**Model-written evaluations** (Anthropic, Dec 2022) use language models to help generate large numbers of test items (e.g. yes/no questions, multi-stage schemas) so labs can discover behaviors faster than pure crowdwork. **HarmBench** (M...

What it is

Model-written evaluations (Anthropic, Dec 2022) use language models to help generate large numbers of test items (e.g. yes/no questions, multi-stage schemas) so labs can discover behaviors faster than pure crowdwork. HarmBench (Mazeika et al., ICML 2024 / PMLR 235) is a standardized evaluation *framework* for automated red teaming and robust refusal—comparing many methods and defenses scientifically.

Why it matters

Human-written suites do not scale to every novel behavior. Automation helps coverage—but inherits model bias and fabrication risk (Anthropic “ouroboros” challenge). HarmBench’s abstract motivation: the field lacked a shared yardstick for automated red-teaming research.

How it works (plain)

  1. Generate candidate tests with models (and/or templates).
  2. Humans verify relevance and labels (Anthropic: crowdworkers often agreed with 90–100% of labels on their generated sets).
  3. Run target models; score behaviors of interest.
  4. For refusal robustness research, use standardized frameworks (HarmBench pointer)—in authorized research settings only.
  5. Feed findings into training/mitigations; re-measure.

Everyday example

A teacher uses an item bank generator for practice quizzes, then still reviews every question before it counts for a grade.

Try it

List three *behaviors* you would want an automated suite to detect (e.g. “reveals secrets from tools,” “agrees with user falsehoods”). Do not design prompts that elicit them.

Myths

⚠️ Myth: Model-written = human-free.
✓ Reality: Anthropic still relies on humans to verify accuracy.
⚠️ Myth: HarmBench is a consumer safety certificate.
✓ Reality: It is a research evaluation framework (paper + open-source repo for researchers)—not a legal compliance badge.
⚠️ Myth: Inverse scaling findings mean bigger is always worse.
✓ Reality: Anthropic reports *some* behaviors worsen with scale or RLHF; treat as measured cases, not universal law.

Sources