Mean, variance, and sampling
The **mean** is a typical value. **Variance** (and standard deviation) describe spread. **Sampling** means learning about a larger world from a smaller set of examples—with uncertainty.
What it is
The mean is a typical value. Variance (and standard deviation) describe spread. Sampling means learning about a larger world from a smaller set of examples—with uncertainty.
Why it matters
Metrics, A/B tests, and “the model scored 92%” all depend on careful sampling language. Overconfident averages mislead product decisions.
How it works (plain)
Averages hide multimodal messes. Spread tells you whether points cluster or scatter. Small samples wobble; report ranges and limitations, not fake certainty.
Everyday example
Class average test score of 80 with everyone near 80 differs from average 80 with scores from 40 to 100.
Try it
Write five numbers. Compute a rough mean and say whether the set feels tight or spread out.
Myths
- ⚠️ Myth: The mean always represents “the typical person.”
- ✓ Reality: Skewed data needs medians and full distributions.
- ⚠️ Myth: Bigger samples erase bias.
- ✓ Reality: Biased sampling stays biased at any size.
Sources
- Khan Academy statistics (verify): https://www.khanacademy.org/math/statistics-probability ↗
- Course 04 probability; Course 18 evaluation
- Elements of AI: https://www.elementsofai.com/ ↗
