COURSE 18L1100% FREE
Verified 2026-08-14

Online evaluation and A/B

Measuring AI changes on live traffic with experiments and guardrails—complementing offline golden sets. Anthropic notes open-ended dialogue A/B tests with humans as a core (if expensive) way to rank helpfulness and harmlessness in realis...

What it is

Measuring AI changes on live traffic with experiments and guardrails—complementing offline golden sets. Anthropic notes open-ended dialogue A/B tests with humans as a core (if expensive) way to rank helpfulness and harmlessness in realistic settings.

MEDIUM PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Online eval: traffic split, guardrail metrics, primary outcome, peeking warning. Tie to causal A/B chapter lightly.

Educational Focus: Complements offline benchmarks.

Why it matters

Offline wins can still hurt latency, cost, or user trust online. NIST AI RMF / GenAI Profile thinking treats measurement as continuous across the lifecycle—not a pre-launch quiz only.

How it works (plain)

Ship to a slice → track task success + guardrails → compare → expand or roll back. Pair with Course 16 A/B literacy. Keep kill switches for safety metrics (escalation rate, policy-flag rate).

Everyday example

A restaurant tries a new recipe on one night’s specials before reprinting the whole menu.

Try it

Name one online guardrail for an AI feature (e.g., escalation rate, refusal quality, citation failure rate).

Myths

⚠️ Myth: Offline eval is enough for LLMs.
✓ Reality: User mix and novel misuse appear live.
⚠️ Myth: Engagement up = success.
✓ Reality: Engagement can rise while task success or safety falls—define gold success first.

Sources