GPU, TPU, Trainium, MI300 landscape
A literacy map of major **AI accelerator families** used in cloud/HPC: NVIDIA **H100** (Hopper), Google **Cloud TPU**, AWS **Trainium** (with Inferentia for serving), and AMD **Instinct MI300** (CDNA 3). Not a buying recommendation—an or...
What it is
A literacy map of major AI accelerator families used in cloud/HPC: NVIDIA H100 (Hopper), Google Cloud TPU, AWS Trainium (with Inferentia for serving), and AMD Instinct MI300 (CDNA 3). Not a buying recommendation—an orientation to primary docs.
Visual Spec & Architecture Diagram
Accelerator landscape comparison cards: GPU / TPU / Trainium / MI300-class — columns: typical use, memory story, ecosystem notes. Plain text, no fake benchmarks as gospel. 'Landscape literacy 2026—verify specs'.
Why it matters
Builders and buyers need to know *what differs*: memory (HBM), matrix units, interconnect, and software stacks—not only brand logos.
How it works (plain)
| Family | One-line idea | Primary doc |
|---|---|---|
| NVIDIA H100 | Hopper Tensor Cores + Transformer Engine; NVLink scale-up | https://www.nvidia.com/en-us/data-center/h100/ ↗ |
| Google TPU | ASIC matrix processor (MXU systolic arrays); Pods/slices/ICI | https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm ↗ |
| AWS Trainium | Purpose-built training chips + Neuron SDK / DLAMI path | https://aws.amazon.com/ai/machine-learning/trainium/getting-started/ ↗ |
| AMD MI300 | CDNA 3 XCDs + HBM3; Infinity Fabric multi-GPU nodes | https://instinct.docs.amd.com/develop/gpu-arch/mi300.html ↗ |
Compare systems with MLPerf when you need shared rules.
Everyday example
Truck vs cargo plane vs freighter—different vehicles for moving “tons,” each with its own ports and fuel story.
Try it
Pick one workload (train vs serve). Write which constraints bite first: memory, interconnect, software lock-in, or $/hour.
Myths
- ⚠️ Myth: One chip generation wins forever.
- ✓ Reality: Stack, price, availability, and model fit change quarterly.
- ⚠️ Myth: Custom ASICs can’t run PyTorch/JAX ecosystems.
- ✓ Reality: Cloud docs emphasize framework support paths (e.g. PyTorch/JAX on TPU; Neuron on Trainium)—verify current support matrices.
