KV cache and inference basics
During autoregressive decoding, models cache key/value tensors (**KV cache**) so they don’t recompute attention over past tokens every step—central to inference speed/memory.
What it is
During autoregressive decoding, models cache key/value tensors (KV cache) so they don’t recompute attention over past tokens every step—central to inference speed/memory.
Why it matters
Explains why long chats get expensive and why batching/serving engineering matters (Course 24).
How it works (plain)
Each new token reuses cached past K/V; memory grows with context length and batch size. Quantization and paging systems manage that growth.
Try it
When a host charges for input tokens on long threads, connect that bill to context + cache reality.
Myths
- ⚠️ Myth: Context is “free” once loaded.
- ✓ Reality: Memory and compute scale with what you keep in play.
Sources
- Course 07 context windows; Course 24 quantization/serving
- Serving system blogs (cite specifically)
