Online evaluation and A/B
Measuring AI changes on live traffic with experiments and guardrails—complementing offline golden sets. Anthropic notes open-ended dialogue A/B tests with humans as a core (if expensive) way to rank helpfulness and harmlessness in realis...
What it is
Measuring AI changes on live traffic with experiments and guardrails—complementing offline golden sets. Anthropic notes open-ended dialogue A/B tests with humans as a core (if expensive) way to rank helpfulness and harmlessness in realistic settings.
Visual Spec & Architecture Diagram
Online eval: traffic split, guardrail metrics, primary outcome, peeking warning. Tie to causal A/B chapter lightly.
Why it matters
Offline wins can still hurt latency, cost, or user trust online. NIST AI RMF / GenAI Profile thinking treats measurement as continuous across the lifecycle—not a pre-launch quiz only.
How it works (plain)
Ship to a slice → track task success + guardrails → compare → expand or roll back. Pair with Course 16 A/B literacy. Keep kill switches for safety metrics (escalation rate, policy-flag rate).
Everyday example
A restaurant tries a new recipe on one night’s specials before reprinting the whole menu.
Try it
Name one online guardrail for an AI feature (e.g., escalation rate, refusal quality, citation failure rate).
Myths
- ⚠️ Myth: Offline eval is enough for LLMs.
- ✓ Reality: User mix and novel misuse appear live.
- ⚠️ Myth: Engagement up = success.
- ✓ Reality: Engagement can rise while task success or safety falls—define gold success first.
Sources
- Anthropic — Challenges in evaluating AI systems: https://www.anthropic.com/research/evaluating-ai-systems ↗
- NIST AI RMF 1.0: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10 ↗
- NIST GenAI Profile: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence ↗
- Course 16 ab-tests; Course 17 monitoring
