Training-serving skew lab
A practice lab to detect when **features differ** between training and live serving—silent accuracy death. Google Cloud’s MLOps level-0 challenges call out handoffs where engineers rebuild features for production APIs, creating skew. Lev...
What it is
A practice lab to detect when features differ between training and live serving—silent accuracy death. Google Cloud’s MLOps level-0 challenges call out handoffs where engineers rebuild features for production APIs, creating skew. Level-1 feature stores exist largely to prevent that.
<!-- IMAGE: same event → two feature paths → diff checker -->
Visual Spec & Architecture Diagram
Skew sources checklist visual: different libraries, missing features online, timezone bugs, training on future labels. Red X / green check examples.
Why it matters
Many “model went bad” stories are skew, not concept drift. Fixing the model while two codepaths disagree wastes weeks.
How it works (plain)
- Pick a sample of real events.
- Compute features two ways (batch/train path vs live/serve path).
- Diff values within tolerance.
- Fail the pipeline if mismatches exceed budget.
- Fix shared code/store—not a one-off notebook patch.
Everyday example
Two cashiers ringing up the same cart with different tax rules. The receipt “looks fine” until audit day.
Try it
Intentionally change timezone handling in one path only; watch diffs light up. Then unify the helper.
Myths
- ⚠️ Myth: Same SQL file guarantees same results everywhere.
- ✓ Reality: Clocks, locales, late data, and caching break parity.
- ⚠️ Myth: Skew only happens at Big Tech scale.
- ✓ Reality: Any train/serve split can diverge.
Sources
- Google Cloud MLOps (skew / feature store): https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning ↗
- Google ML Crash Course (skew notes): https://developers.google.com/machine-learning/crash-course ↗
- Course 17 feature stores; Course 02 leakage
