Quantization and serving basics
**Quantization** stores/computes model weights (and sometimes activations) in fewer bits to reduce memory and often cost/latency. **Serving** is running models reliably for live traffic. Modern accelerators advertise low-precision matrix...
What it is
Quantization stores/computes model weights (and sometimes activations) in fewer bits to reduce memory and often cost/latency. Serving is running models reliably for live traffic. Modern accelerators advertise low-precision matrix paths (e.g. FP8 on Hopper Transformer Engine; INT8/FP8 matrix throughput on MI300 CDNA 3 docs).
Visual Spec & Architecture Diagram
Quantization bit-width ladder: FP32 → FP16/BF16 → INT8 → INT4 with size/speed/quality tradeoff icons; serving batching sketch.
Why it matters
Many “we can’t afford that model” problems are serving/encoding problems as much as algorithm problems. Inference energy and data-center load (IEA/LBNL) scale with how inefficiently you serve.
How it works (plain)
Shrink numeric precision carefully → measure quality drop on your evals → batch requests → cache where safe → monitor tail latency. Not all tasks tolerate aggressive quantization. MLPerf Inference scenarios (offline, server, single/multi-stream) exist to compare serving regimes fairly.
Everyday example
Streaming music at a lower bitrate—smaller/faster, sometimes audible loss.
Try it
Ask whether your provider’s “cheap” tier is a smaller model, quantized model, or both.
Myths
- ⚠️ Myth: Quantization is free quality.
- ✓ Reality: Always re-eval on your tasks.
- ⚠️ Myth: Peak chip TFLOPS predict serving cost.
- ✓ Reality: Utilization, batching, and memory bandwidth dominate.
Sources
- MLPerf Inference docs: https://docs.mlcommons.org/inference/index_gh/ ↗
- NVIDIA H100 (Transformer Engine / FP8): https://www.nvidia.com/en-us/data-center/h100/ ↗
- AMD MI300 peak table (FP8/INT8 matrix): https://instinct.docs.amd.com/develop/gpu-arch/mi300.html ↗
