Sampling: temperature and top-p
After a model scores possible next tokens, **decoding** chooses among them. **Temperature** and **top-p** (nucleus sampling) are common knobs that trade boredom for creativity—and sometimes for chaos.
What it is
After a model scores possible next tokens, decoding chooses among them. Temperature and top-p (nucleus sampling) are common knobs that trade boredom for creativity—and sometimes for chaos.
Why it matters
Support bots often want low randomness. Brainstorming tools want more. The same model can feel “strict” or “wild” from decoding alone.
How it works (plain)
- Lower temperature → peakier distribution → more predictable
- Higher temperature → flatter distribution → more variety
- Top-p keeps only the smallest set of tokens whose probabilities sum to \(p\), then samples inside that set
Neither knob installs truthfulness.
Everyday example
Temperature near 0 for extracting fields from a form. Higher temperature for slogan ideas—then you filter.
Try it
Ask for five product names twice: once “be deterministic,” once “be inventive.” Compare overlap.
Myths
- ⚠️ Myth: Temperature 0 guarantees correctness.
- ✓ Reality: It reduces sampling noise; wrong high-probability answers remain.
- ⚠️ Myth: Higher creativity settings mean deeper intelligence.
- ✓ Reality: They change randomness, not understanding.
Sources
- Holtzman et al., “The Curious Case of Neural Text Degeneration” (nucleus sampling): https://arxiv.org/abs/1904.09751 ↗
- Hugging Face generation docs: https://huggingface.co/docs/transformers/main/en/generation_strategies ↗
- Provider decoding docs (cite the model you use)
