COURSE 07L1100% FREE
Verified 2026-08-10

Speculative decoding overview

A serving trick: a **smaller draft model** proposes several tokens; the **large model** checks them in parallel and accepts a prefix—speeding inference without changing the big model’s distribution (when done correctly).

What it is

A serving trick: a smaller draft model proposes several tokens; the large model checks them in parallel and accepts a prefix—speeding inference without changing the big model’s distribution (when done correctly).

Why it matters

Latency and cost dominate product experience. Speculative decoding is one reason some APIs feel snappier.

How it works (plain)

Draft fast → verify with the big model → keep matching tokens → resume. Users still see the large model’s answers; systems engineers see throughput gains.

Everyday example

An apprentice types ahead; the expert accepts until the first wrong keystroke—then takes over.

Try it

When a vendor claims speedups, ask whether quality/distribution is preserved and on which hardware.

Myths

⚠️ Myth: Speculative decoding changes what the model “knows.”
✓ Reality: It’s mainly a systems acceleration (with implementation caveats).
⚠️ Myth: It always helps.
✓ Reality: Gains depend on acceptance rates and workload shape.

Sources

  • Course 07 sampling; Course 24 compute
  • Leviathan et al. / related speculative decoding papers (cite specifically)