Feature stores
A **feature store** is shared infrastructure for defining, serving, and documenting ML features consistently for training and live inference. Google Cloud’s MLOps level-1 guidance lists a feature store as an optional but powerful compone...
What it is
A feature store is shared infrastructure for defining, serving, and documenting ML features consistently for training and live inference. Google Cloud’s MLOps level-1 guidance lists a feature store as an optional but powerful component for continuous training: one definition, batch history for training, low-latency values for serving.
<!-- IMAGE: offline store ↔ online store with shared feature definitions -->
Visual Spec & Architecture Diagram
Feature store: Offline store (batch, training) vs Online store (low-latency serving) with same feature definitions; point-in-time correct joins callout; training-serving consistency arrow.
Why it matters
Training-serving skew—different feature logic in notebooks vs production—is a classic silent killer. Feature stores attack that by making the same feature code/metadata the source for experiments, CT pipelines, and online prediction.
How it works (plain)
- Define features once (name, owner, transform, freshness SLA).
- Backfill historical values for training (point-in-time correct).
- Serve low-latency values online for the same entity keys.
- Monitor freshness and drift.
- Document owners so on-call knows who to page.
Everyday example
A shared spice pantry with labeled jars vs each cook inventing “mild chili” differently every night.
Try it
List 5 features your product needs and who owns their definition today. Mark any that are computed differently in training vs serving.
Myths
- ⚠️ Myth: A feature store replaces data engineering.
- ✓ Reality: It organizes features; pipelines and quality still matter.
- ⚠️ Myth: Only Big Tech needs this.
- ✓ Reality: Even small teams need *consistency*—sometimes a lightweight store or a strict shared library.
- ⚠️ Myth: Online cache = correct training history.
- ✓ Reality: Without point-in-time joins, you leak the future into labels.
Sources
- Google Cloud MLOps (feature store in level 1): https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning ↗
- Course 03 feature engineering; Course 02 leakage; Course 17 skew lab
- TFX / pipeline platforms that materialize transforms: https://www.tensorflow.org/tfx ↗
