Positional encodings
Transformers treat tokens as a set unless you tell them **order**. **Positional encodings** (or positional embeddings) inject “this is word 1, this is word 2…” so grammar and sequence make sense.
What it is
Transformers treat tokens as a set unless you tell them order. Positional encodings (or positional embeddings) inject “this is word 1, this is word 2…” so grammar and sequence make sense.
Why it matters
Without position, “dog bites man” and “man bites dog” look too similar to the model’s bag of tokens. Position is how order enters the architecture.
How it works (plain)
Early transformers added fixed sine/cosine patterns. Many modern models learn position vectors or use relative position schemes so long contexts behave better.
Everyday example
Sheet music without bar numbers—or a script with lines shuffled—loses meaning. Position restores sequence.
Try it
Scramble a sentence’s words and ask what broke (subject/object, tense markers). That’s why position matters.
Myths
- ⚠️ Myth: Attention alone knows order.
- ✓ Reality: Plain attention is permutation-sensitive only after positions are added (or built into the attention bias).
- ⚠️ Myth: One positional scheme fits every context length.
- ✓ Reality: Long-context models often change position methods; check model cards.
Sources
- Vaswani et al., 2017: https://arxiv.org/abs/1706.03762 ↗
- Course 07 transformers; self-attention; context-windows
