Overfitting and generalization
**Overfitting** means a model looks great on the training examples but fails on new ones—it memorized quirks instead of learning the pattern. **Generalization** means performance holds up on fresh data from the same kind of situation.
What it is
Overfitting means a model looks great on the training examples but fails on new ones—it memorized quirks instead of learning the pattern.
Generalization means performance holds up on fresh data from the same kind of situation.
Why it matters
A model that aces yesterday’s homework and fails tomorrow’s quiz is useless in production—and dangerous if people trust the homework score.
How it works (plain)
Think of a student who memorizes answer keys vs one who learns the method. Held-out validation/test sets are the surprise quiz.
Common defenses: more data, simpler models, regularization, early stopping, data augmentation, cross-validation.
Everyday example
A fraud model trained only on last month’s scams may miss next month’s new scam style (distribution shift)—a cousin of generalization failure.
Try it
Fit a story to 5 points perfectly in your head with a wild explanation, then ask if a 6th point would fit. That discomfort is overfitting intuition.
Myths
- ⚠️ Myth: Bigger models always generalize better.
- ✓ Reality: Capacity helps *and* enables memorization; data, regularization, and evaluation decide the outcome.
- ⚠️ Myth: Perfect train accuracy is the goal.
- ✓ Reality: Useful test performance is the goal.
Sources
- Google ML Crash Course (generalization): https://developers.google.com/machine-learning/crash-course ↗
- ANN live: https://www.ainerdnetwork.com/learn/overfitting-and-generalization ↗
