MLOps overview
**MLOps** is the practice of treating ML like a production system: data, training, deployment, monitoring, and rollback—not a one-off notebook. Google Cloud’s architecture guide defines MLOps as unifying ML *development* and *operations*...
What it is
MLOps is the practice of treating ML like a production system: data, training, deployment, monitoring, and rollback—not a one-off notebook. Google Cloud’s architecture guide defines MLOps as unifying ML *development* and *operations* with automation and monitoring at every step (integration, testing, release, deployment, infrastructure).
<!-- IMAGE: notebook → pipeline → serve → monitor loop -->
Visual Spec & Architecture Diagram
MLOps loop: Data → Train → Validate → Deploy → Monitor → Retrain, with Model Registry and Feature Store as side hubs. Title: 'MLOps lifecycle'.
Why it matters
A strong offline metric is not a production system. Google’s framing (adapting Sculley et al.) shows ML *code* is a small box surrounded by configuration, data collection/verification, testing, serving, and monitoring. Most failures live in that surrounding system.
How it works (plain)
- Version data, code, and configs together.
- Train in a reproducible pipeline (not only an interactive notebook).
- Validate data and model before promote.
- Deploy with a rollback path.
- Monitor quality, latency, cost, and safety signals.
- Retrain or roll back when live behavior drifts.
LLM apps add prompt/versioning, tool-permission reviews, and eval gates (Course 07/08/10).
Everyday example
A restaurant that scales a recipe needs suppliers, checklists, and health inspections—not only a tasty first plate. MLOps is the kitchen ops for models.
Try it
Pick one AI feature at work. Write the top 3 things you would monitor in the first 24 hours after a change (quality, latency, cost, or safety).
Myths
- ⚠️ Myth: Test-set accuracy equals production readiness.
- ✓ Reality: Drift, skew, latency, and abuse appear only live.
- ⚠️ Myth: MLOps is only for huge teams.
- ✓ Reality: Solo builders still need versioning, a registry habit, and basic monitoring.
- ⚠️ Myth: MLOps = “deploy the model API.”
- ✓ Reality: Mature setups deploy *pipelines* that can retrain and re-serve (Google Cloud levels 0→2).
Sources
- Google Cloud — MLOps continuous delivery & automation: https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning ↗
- AWS — Planning for successful MLOps: https://docs.aws.amazon.com/prescriptive-guidance/latest/ml-operations-planning/introduction.html ↗
- Microsoft Learn — Operationalize ML models (MLOps): https://learn.microsoft.com/en-us/training/paths/build-first-machine-operations-workflow/ ↗
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
