COURSE 07L1100% FREE
Verified 2026-08-10

Query, key, value attention

Inside transformers, **attention** uses three roles for each token: a **query** (“what am I looking for?”), **keys** (“what do I contain?”), and **values** (“what content do I pass along?”).

What it is

Inside transformers, attention uses three roles for each token: a query (“what am I looking for?”), keys (“what do I contain?”), and values (“what content do I pass along?”).

Why it matters

This is the mechanism behind “the model focused on that word.” Understanding Q/K/V demystifies demos of attention maps—without claiming they are full explanations of meaning.

How it works (plain)

  1. Each token makes a query, key, and value (via learned projections)
  2. Queries are compared to keys to get match scores
  3. Scores become weights
  4. Weighted values are combined into a new representation

Multi-head attention repeats this in parallel “heads” with different projections.

Everyday example

In a library, your question (query) matches catalog cards (keys); you pull the books’ contents (values) that matched best.

Try it

Read one sentence and mark which earlier words a later word must “look at” to make sense (pronouns are a good start).

Myths

⚠️ Myth: Bright attention weights prove the model “knows” causation.
✓ Reality: Weights are useful clues, not complete interpretability.
⚠️ Myth: One head does all the work.
✓ Reality: Different heads often specialize; still not magic.

Sources