Synthetic data basics
**Synthetic data** is generated by simulations or models rather than collected directly from the world. It can help privacy and coverage—or amplify junk.
What it is
Synthetic data is generated by simulations or models rather than collected directly from the world. It can help privacy and coverage—or amplify junk.
Why it matters
Teams use synthetic data to fill rare classes and protect PII. Without checks, models train on their own fantasies.
How it works (plain)
Start from real distributions or simulators → generate → validate against holdout real data → mix carefully. Never treat synthetic-only metrics as production proof.
Everyday example
A driving simulator creates rare storm scenarios you can’t safely film—still need real-road validation.
Try it
List one place synthetic data would help your domain and one way it could mislead.
Myths
- ⚠️ Myth: Synthetic data is automatically unbiased.
- ✓ Reality: It inherits generator and design biases.
- ⚠️ Myth: More synthetic always helps.
- ✓ Reality: Model collapse / distribution drift risks rise when synthetic dominates.
Sources
- Course 02 splits; Course 11 generative media
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
