Perplexity intuition
**Perplexity** summarizes how surprised a language model is by a text set—lower usually means better next-token prediction on that set.
What it is
Perplexity summarizes how surprised a language model is by a text set—lower usually means better next-token prediction on that set.
Why it matters
Common training/eval metric for LMs—but not a direct measure of helpfulness, truth, or safety.
How it works (plain)
If a model assigns higher probability to the true next tokens, perplexity drops. Domain mismatch raises it. Don’t confuse leaderboard perplexity with product quality.
Everyday example
A spellchecker “surprised” by medical jargon it never saw.
Try it
When a paper cites perplexity, ask: on which data, compared to what?
Myths
- ⚠️ Myth: Lowest perplexity model is the best assistant.
- ✓ Reality: Alignment, tools, and eval tasks matter more for products.
Sources
- Course 07 language-models; Course 18 evaluation
- Jurafsky & Martin: https://web.stanford.edu/~jurafsky/slp3/ ↗
