Safety eval suites
Bundles of tests for disallowed behaviors, privacy leaks, and tool misuse—run like a regression suite. Distinct from capability leaderboards: here the goal is **robust refusal / safe behavior**, not higher quiz scores.
What it is
Bundles of tests for disallowed behaviors, privacy leaks, and tool misuse—run like a regression suite. Distinct from capability leaderboards: here the goal is robust refusal / safe behavior, not higher quiz scores.
Visual Spec & Architecture Diagram
Safety suite layers: automated classifiers → curated prompts → human red team → production monitors. Coverage heatmap fake.
Why it matters
Capability evals alone miss harm pathways. HarmBench (ICML 2024) exists because automated red-teaming lacked a standardized measurement frame for attacks *and* defenses—literacy: the field needs comparable refusal metrics, not ad-hoc demos.
How it works (plain)
Policy → test cases by category → automated + expert review → severity scores → block ship on high severity → retest after fixes. No public exploit cookbooks in this curriculum.
UK AISI frames safety-relevant capability measurement as early-warning evidence for policymakers—not a regulator’s “safe/unsafe” stamp.
Everyday example
Aviation checklists: you re-run critical items after any change to the aircraft—not once at delivery.
Try it
Write 5 policy-breaking *goals* (not attack steps) for your product and map each to a severity (low/med/high/critical).
Myths
- ⚠️ Myth: One safety suite forever.
- ✓ Reality: New tools/features need new cases.
- ⚠️ Myth: Higher MMLU implies safer.
- ✓ Reality: Orthogonal axes—measure both.
Sources
- HarmBench (PMLR): https://proceedings.mlr.press/v235/mazeika24a ↗
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
- UK AISI approach: https://www.aisi.gov.uk/blog/our-approach-to-evaluations ↗
- OWASP LLM Top 10: https://owasp.org/www-project-top-10-for-large-language-model-applications/ ↗
- NIST GenAI Profile: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence ↗
