COURSE 07L1100% FREE
Verified 2026-08-10

KV cache and inference basics

During autoregressive decoding, models cache key/value tensors (**KV cache**) so they don’t recompute attention over past tokens every step—central to inference speed/memory.

What it is

During autoregressive decoding, models cache key/value tensors (KV cache) so they don’t recompute attention over past tokens every step—central to inference speed/memory.

Why it matters

Explains why long chats get expensive and why batching/serving engineering matters (Course 24).

How it works (plain)

Each new token reuses cached past K/V; memory grows with context length and batch size. Quantization and paging systems manage that growth.

Try it

When a host charges for input tokens on long threads, connect that bill to context + cache reality.

Myths

⚠️ Myth: Context is “free” once loaded.
✓ Reality: Memory and compute scale with what you keep in play.

Sources

  • Course 07 context windows; Course 24 quantization/serving
  • Serving system blogs (cite specifically)