Ensembles and boosting (deeper)
**Ensembles** combine multiple models. **Boosting** builds models in sequence, each focusing more on previous mistakes (e.g., gradient boosting families).
What it is
Ensembles combine multiple models. Boosting builds models in sequence, each focusing more on previous mistakes (e.g., gradient boosting families).
Why it matters
Tabular ML in industry still often wins with gradient boosting. Knowing why prevents “deep learning everything” mistakes.
How it works (plain)
Bagging averages parallel diverse models (random forests). Boosting adds specialists for residuals. Diversity + aggregation reduce error—until you overfit or leak.
Everyday example
A panel of reviewers catches more issues than one reviewer—if they aren’t clones.
Try it
Compare a single decision stump story vs a sequence of stump corrections on a tiny example.
Myths
- ⚠️ Myth: Neural nets always beat boosting on spreadsheets.
- ✓ Reality: For many tabular tasks, boosting remains extremely strong.
- ⚠️ Myth: More trees always help.
- ✓ Reality: Diminishing returns + overfit—use validation.
Sources
- Course 03 classical-ml-models
- scikit-learn ensemble docs: https://scikit-learn.org/ ↗
- Friedman gradient boosting literature (cite when teaching)
