Model-written evals and HarmBench
**Model-written evaluations** (Anthropic, Dec 2022) use language models to help generate large numbers of test items (e.g. yes/no questions, multi-stage schemas) so labs can discover behaviors faster than pure crowdwork. **HarmBench** (M...
What it is
Model-written evaluations (Anthropic, Dec 2022) use language models to help generate large numbers of test items (e.g. yes/no questions, multi-stage schemas) so labs can discover behaviors faster than pure crowdwork. HarmBench (Mazeika et al., ICML 2024 / PMLR 235) is a standardized evaluation *framework* for automated red teaming and robust refusal—comparing many methods and defenses scientifically.
Why it matters
Human-written suites do not scale to every novel behavior. Automation helps coverage—but inherits model bias and fabrication risk (Anthropic “ouroboros” challenge). HarmBench’s abstract motivation: the field lacked a shared yardstick for automated red-teaming research.
How it works (plain)
- Generate candidate tests with models (and/or templates).
- Humans verify relevance and labels (Anthropic: crowdworkers often agreed with 90–100% of labels on their generated sets).
- Run target models; score behaviors of interest.
- For refusal robustness research, use standardized frameworks (HarmBench pointer)—in authorized research settings only.
- Feed findings into training/mitigations; re-measure.
Everyday example
A teacher uses an item bank generator for practice quizzes, then still reviews every question before it counts for a grade.
Try it
List three *behaviors* you would want an automated suite to detect (e.g. “reveals secrets from tools,” “agrees with user falsehoods”). Do not design prompts that elicit them.
Myths
- ⚠️ Myth: Model-written = human-free.
- ✓ Reality: Anthropic still relies on humans to verify accuracy.
- ⚠️ Myth: HarmBench is a consumer safety certificate.
- ✓ Reality: It is a research evaluation framework (paper + open-source repo for researchers)—not a legal compliance badge.
- ⚠️ Myth: Inverse scaling findings mean bigger is always worse.
- ✓ Reality: Anthropic reports *some* behaviors worsen with scale or RLHF; treat as measured cases, not universal law.
Sources
- Anthropic — Model-Written Evaluations: https://www.anthropic.com/research/discovering-language-model-behaviors-with-model-written-evaluations ↗
- HarmBench PMLR page: https://proceedings.mlr.press/v235/mazeika24a ↗
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
