PII in training data
Personally identifiable information accidentally (or carelessly) ending up in training sets, prompts, or logs—and how to minimize it.
What it is
Personally identifiable information accidentally (or carelessly) ending up in training sets, prompts, or logs—and how to minimize it.
Why it matters
Privacy incidents destroy trust. Connects Course 02 provenance to Course 19 privacy.
How it works (plain)
Inventory fields → redact/minimize → access control → retention limits → never paste secrets into public models. Prefer synthetic or aggregated data when possible.
Try it
Scan one CSV header list; mark PII vs safe features.
Myths
- ⚠️ Myth: “We anonymized names” always equals safe.
- ✓ Reality: Quasi-identifiers still re-identify.
Sources
- Course 19 privacy; Course 02 consent; regulator guidance (primary)
