Query, key, value attention
Inside transformers, **attention** uses three roles for each token: a **query** (“what am I looking for?”), **keys** (“what do I contain?”), and **values** (“what content do I pass along?”).
What it is
Inside transformers, attention uses three roles for each token: a query (“what am I looking for?”), keys (“what do I contain?”), and values (“what content do I pass along?”).
Why it matters
This is the mechanism behind “the model focused on that word.” Understanding Q/K/V demystifies demos of attention maps—without claiming they are full explanations of meaning.
How it works (plain)
- Each token makes a query, key, and value (via learned projections)
- Queries are compared to keys to get match scores
- Scores become weights
- Weighted values are combined into a new representation
Multi-head attention repeats this in parallel “heads” with different projections.
Everyday example
In a library, your question (query) matches catalog cards (keys); you pull the books’ contents (values) that matched best.
Try it
Read one sentence and mark which earlier words a later word must “look at” to make sense (pronouns are a good start).
Myths
- ⚠️ Myth: Bright attention weights prove the model “knows” causation.
- ✓ Reality: Weights are useful clues, not complete interpretability.
- ⚠️ Myth: One head does all the work.
- ✓ Reality: Different heads often specialize; still not magic.
Sources
- “Attention Is All You Need” (Vaswani et al., 2017) — https://arxiv.org/abs/1706.03762 ↗
- Course 07 self-attention; transformers
- ANN Learn transformers pages
