Tokenization
Models do not read English letters the way you do. **Tokenization** splits text into **tokens**—chunks like whole words, word pieces, or punctuation—that the model scores.
What it is
Models do not read English letters the way you do. Tokenization splits text into tokens—chunks like whole words, word pieces, or punctuation—that the model scores.
Why it matters
Tokens affect cost (you pay per token), length limits (context windows count tokens), and weird failures (languages or code that tokenize poorly).
How it works (plain)
A tokenizer uses a fixed vocabulary learned from lots of text. Common words may be one token; rare words may split into pieces. Spaces and punctuation matter.
Everyday example
The snarky demo: some models “struggle” with letter counting because they see tokens, not characters. The deeper lesson: the model’s world is token sequences.
Try it
Paste a sentence into a public tokenizer visualizer from a model provider’s docs (when available) and count tokens vs words.
Myths
- ⚠️ Myth: One token equals one word.
- ✓ Reality: Often word pieces; multilingual text varies a lot.
- ⚠️ Myth: Tokenization is just a detail.
- ✓ Reality: It shapes pricing, context, and some failure modes.
Sources
- Hugging Face tokenizers docs: https://huggingface.co/docs/transformers/tokenizer_summary ↗
- OpenAI tokenizer explainer (educational): https://platform.openai.com/tokenizer ↗
- ANN live: https://www.ainerdnetwork.com/learn/tokenization ↗
