Imbalanced classes
When some labels are rare—fraud, defects, disease flags—**class imbalance** makes naive training and accuracy metrics lie.
What it is
When some labels are rare—fraud, defects, disease flags—class imbalance makes naive training and accuracy metrics lie.
Why it matters
Many valuable problems are imbalanced. Fixing them is mostly careful metrics, sampling, and costs—not magic architectures.
How it works (plain)
Options: better metrics, class weights, resampling, threshold tuning, gathering more rare examples, anomaly approaches. Always validate on realistic prevalence.
Everyday example
Predicting “will this package be damaged?” when damage is 0.5% of rows.
Try it
Compute majority-class baseline accuracy for a 1% positive problem (hint: 99%).
Myths
- ⚠️ Myth: SMOTE fixes everything.
- ✓ Reality: Synthetic minority samples can distort; verify on real holdouts.
- ⚠️ Myth: You must balance classes to 50/50 always.
- ✓ Reality: Match the decision costs and prevalence you will see.
Sources
- Course 03 metrics; Course 02 synthetic data
- scikit-learn imbalanced handling guides: https://scikit-learn.org/ ↗
