Self-attention
**Self-attention** lets each token in a sequence build a weighted mix of information from other tokens in that same sequence. “Self” means the sequence attends to itself.
What it is
Self-attention lets each token in a sequence build a weighted mix of information from other tokens in that same sequence. “Self” means the sequence attends to itself.
Why it matters
This is the mechanism people point to when they say transformers “relate words across a sentence.” Pronouns, long dependencies, and code structure all benefit.
How it works (plain)
For each token, ask: “Given what I am trying to predict, which other tokens should I look at, and how much?” The model learns those looking patterns from data—not from a human writing grammar rules.
Everyday example
In “The trophy did not fit in the suitcase because it was too big,” attention patterns help the model decide whether “it” is more about trophy or suitcase—sometimes correctly, sometimes not.
Try it
Underline a pronoun in a sentence and arrow to candidate nouns. That manual disambiguation is the human version of “where to attend.”
Myths
- ⚠️ Myth: Attention weights are a full explanation of reasoning.
- ✓ Reality: They are a useful clue, not a complete interpretability proof.
- ⚠️ Myth: More attention heads always mean smarter models.
- ✓ Reality: Heads help capacity; quality depends on training and data.
Sources
- Vaswani et al., 2017: https://arxiv.org/abs/1706.03762 ↗
- Illustrated Transformer: https://jalammar.github.io/illustrated-transformer/ ↗
- ANN live: https://www.ainerdnetwork.com/learn/self-attention ↗
