Monitoring and drift
Watching live AI systems for **data drift**, **performance drift**, cost spikes, and safety incidents—then alerting owners with a runbook. Google Cloud’s MLOps guide treats monitoring as the stage that closes the loop: live stats can tri...
What it is
Watching live AI systems for data drift, performance drift, cost spikes, and safety incidents—then alerting owners with a runbook. Google Cloud’s MLOps guide treats monitoring as the stage that closes the loop: live stats can trigger retraining or a new experiment cycle. AWS MLOps planning likewise treats monitoring as a first-class lifecycle area beside data, training, and deployment.
<!-- IMAGE: baseline distribution vs live distribution with alert threshold -->
Visual Spec & Architecture Diagram
Drift charts: overlay of training vs production feature distributions (PSI-style); prediction drift over calendar; label spike 'data drift' vs 'concept drift' definitions.
Why it matters
Models expire quietly. Offline metrics stay frozen while the world moves. Monitoring is how MLOps earns its keep—and how you notice silent failures before customers do.
How it works (plain)
- Track input stats and output distributions vs a baseline window.
- Track task success proxies (conversion, override rate, ticket rate)—not only model loss.
- Track latency, errors, and cost.
- Tie alerts to owners and runbooks (page, roll back, or retrain).
- Prefer canary/shadow traffic before full promote.
Everyday example
A smoke alarm is useless without a fire exit plan. Metrics without owners are noise.
Try it
Pick one metric you’d alert on within 24 hours of a bad deploy. Write the person/role who owns the response.
Myths
- ⚠️ Myth: Last month’s batch accuracy equals health.
- ✓ Reality: Live traffic shifts.
- ⚠️ Myth: More alerts = safer.
- ✓ Reality: Alert fatigue hides real fires.
- ⚠️ Myth: Monitoring averages is enough.
- ✓ Reality: Slice failures (region, device, language) hide in averages.
Sources
- Google Cloud MLOps: https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning ↗
- AWS MLOps planning (monitoring area): https://docs.aws.amazon.com/prescriptive-guidance/latest/ml-operations-planning/introduction.html ↗
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
- Course 18 online eval
