COURSE 24L1100% FREE
Verified 2026-08-14

Quantization and serving basics

**Quantization** stores/computes model weights (and sometimes activations) in fewer bits to reduce memory and often cost/latency. **Serving** is running models reliably for live traffic. Modern accelerators advertise low-precision matrix...

What it is

Quantization stores/computes model weights (and sometimes activations) in fewer bits to reduce memory and often cost/latency. Serving is running models reliably for live traffic. Modern accelerators advertise low-precision matrix paths (e.g. FP8 on Hopper Transformer Engine; INT8/FP8 matrix throughput on MI300 CDNA 3 docs).

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Quantization bit-width ladder: FP32 → FP16/BF16 → INT8 → INT4 with size/speed/quality tradeoff icons; serving batching sketch.

Educational Focus: Makes compression tradeoffs tangible.

Why it matters

Many “we can’t afford that model” problems are serving/encoding problems as much as algorithm problems. Inference energy and data-center load (IEA/LBNL) scale with how inefficiently you serve.

How it works (plain)

Shrink numeric precision carefully → measure quality drop on your evals → batch requests → cache where safe → monitor tail latency. Not all tasks tolerate aggressive quantization. MLPerf Inference scenarios (offline, server, single/multi-stream) exist to compare serving regimes fairly.

Everyday example

Streaming music at a lower bitrate—smaller/faster, sometimes audible loss.

Try it

Ask whether your provider’s “cheap” tier is a smaller model, quantized model, or both.

Myths

⚠️ Myth: Quantization is free quality.
✓ Reality: Always re-eval on your tasks.
⚠️ Myth: Peak chip TFLOPS predict serving cost.
✓ Reality: Utilization, batching, and memory bandwidth dominate.

Sources