COURSE 16L2100% FREE
Verified 2026-08-14

Identification and confounding

**Identification** asks: *given assumptions*, can the causal effect we care about be written in terms of observable distributions? **Estimation** asks how to compute that from finite data. DoWhy’s docs insist on separating those steps. C...

What it is

Identification asks: *given assumptions*, can the causal effect we care about be written in terms of observable distributions? Estimation asks how to compute that from finite data. DoWhy’s docs insist on separating those steps. Confounding is the usual reason identification fails for naive comparisons.

Why it matters

ML can estimate almost any conditional mean. That does not mean the conditional mean equals an interventional effect. UCLA: experiments are one deconfounding tool—not the only one—but without *some* deconfounding strategy, action advice is guesswork.

How it works (plain)

  1. Name the causal question (the intervention).
  2. State assumptions (DAG and/or potential-outcomes conditions).
  3. Check identification (backdoor set? instrument? experiment?).
  4. Only then pick an estimator (regression, IPW, DR, DML, …).
  5. Stress-test assumptions (sensitivity, refutation).

Hernán & Robins walk this ladder from “without models” through longitudinal g-methods.

Everyday example

Knowing ice cream sales predict drownings does not identify the effect of banning ice cream—summer confounds; the interventional query isn’t identified from that association alone.

Try it

For “discount → revenue,” list one randomized design that identifies the effect and one observational story that might *not* (who self-selects into discounts?).

Myths

⚠️ Myth: Identification is a vibe check on model accuracy.
✓ Reality: It’s a logic step about assumptions + data availability (DoWhy; Pearl calculus).
⚠️ Myth: If the pipeline runs, the effect is identified.
✓ Reality: Software will happily estimate undefined targets if you skip the ID step.

Sources