COURSE 24L1100% FREE
Verified 2026-08-14

GPUs, TPUs, and training compute

**GPUs** and **TPUs** (and similar accelerators) speed the matrix math behind training and inference. **Training compute** is the total work spent to train a model—not the same as “how big it looks on a slide.” Google designed Cloud TPUs...

What it is

GPUs and TPUs (and similar accelerators) speed the matrix math behind training and inference. Training compute is the total work spent to train a model—not the same as “how big it looks on a slide.” Google designed Cloud TPUs as matrix processors (systolic arrays / MXUs) specialized for neural nets—not general desktop CPUs.

Why it matters

Hardware access shapes who can train large models and what inference costs in production. ANN live page preserved via this slug. Cross-vendor literacy: NVIDIA Hopper H100 (Transformer Engine, HBM), Google TPU VMs, AWS Trainium, AMD Instinct MI300 (CDNA 3).

How it works (plain)

Training repeatedly pushes huge batches of math through accelerators. Inference uses them to answer live traffic—often the bigger bill at scale. Memory size and bandwidth can bottleneck before raw FLOPs marketing does. Multi-chip training needs fast interconnect (NVLink, TPU ICI, Infinity Fabric).

Everyday example

A bakery’s industrial oven vs a home oven—same recipes conceptually, different throughput and energy.

Try it

For one AI product you pay for, ask whether pricing tracks tokens, minutes of GPU, or seats.

Myths

⚠️ Myth: Any GPU makes any model fast.
✓ Reality: VRAM/HBM, interconnect, and software stacks (CUDA, Neuron, ROCm/XLA) matter.
⚠️ Myth: Training compute equals intelligence.
✓ Reality: Data quality, algorithms, and eval design matter too.

Sources