A/B tests as interventions
Randomized experiments that **intervene** on what users see—so differences in outcomes can support causal claims (with caveats). UCLA’s DAG materials call well-designed experiments / RCTs a gold standard because randomization balances co...
What it is
Randomized experiments that intervene on what users see—so differences in outcomes can support causal claims (with caveats). UCLA’s DAG materials call well-designed experiments / RCTs a gold standard because randomization balances confounders across groups. Stanford MS&E 226 teaches randomized experiments inside its causality unit (Rubin causal model / potential outcomes).
Visual Spec & Architecture Diagram
A/B as intervention: users randomly assigned to Treatment/Control (do(T)); outcome Y measured; arrow 'randomization breaks confounding paths'. Contrast observational click data with self-selection.
Why it matters
Observational ML metrics ≠ “if we ship this, metrics move.” A/B tests are the product workhorse intervention. EconML’s public use cases also show imperfect compliance and nudges—experiments aren’t always simple two-cell switches.
How it works (plain)
Randomize assignment → measure outcomes → compare with uncertainty → watch peeking and multiple testing. Ethics: don’t experiment irresponsibly on vulnerable users. When you can’t force treatment (e.g. membership signup), instrumental-variable style analyses appear (EconML TripAdvisor-style DRIV example on Microsoft Research page).
Everyday example
Showing half of visitors a new checkout button and comparing purchase rate—assignment is the intervention.
Try it
Design an A/B for an AI feature: primary metric, guardrail metric, stop rule, and one segment you’ll check for harm.
Myths
- ⚠️ Myth: Significant p-value means ship forever.
- ✓ Reality: Practical significance, segments, and long-term effects matter (MS&E 226 also separates prediction/inference/causality mindsets).
- ⚠️ Myth: If you ran an A/B, observational confounding is impossible forever after.
- ✓ Reality: After ship, feedback loops and selection return—keep causal discipline.
Sources
- UCLA DAG intro (experiments & randomization): https://stats.oarc.ucla.edu/wp-content/uploads/2025/06/Intro-do-DAGs-UCLA-OARC.html ↗
- Stanford MS&E 226: https://web.stanford.edu/class/msande226/ ↗
- Hernán & Robins *What If*: https://miguelhernan.org/whatifbook ↗
- EconML use cases: https://www.microsoft.com/en-us/research/project/econml/ ↗
