Networking for distributed training
Why interconnect bandwidth/latency matter when training across many accelerators—and how weak networks erase GPU spend. Vendor stacks emphasize dedicated links: NVIDIA NVLink, AMD Infinity Fabric between MI300X OAMs, Google TPU inter-chi...
What it is
Why interconnect bandwidth/latency matter when training across many accelerators—and how weak networks erase GPU spend. Vendor stacks emphasize dedicated links: NVIDIA NVLink, AMD Infinity Fabric between MI300X OAMs, Google TPU inter-chip interconnect (ICI) within slices (and DCN for multislice).
Why it matters
Buying GPUs without networking is a classic waste pattern. Large LLM training jobs are often communication-bound at scale.
How it works (plain)
Gradients/parameters must sync. Slow links → accelerators wait. Topology and collective ops dominate large jobs. Cloud quotes that omit interconnect class are incomplete.
Everyday example
Eight chefs with one tiny pass-through window—more cooks don’t equal more plates.
Try it
Ask a cloud quote for multi-GPU: what’s the interconnect story (NVLink/ICI/Fabric class vs generic Ethernet only)?
Myths
- ⚠️ Myth: 8× GPUs means 8× speed if Ethernet is “fine.”
- ✓ Reality: Collectives can bottleneck hard.
Sources
- Google TPU system architecture (ICI / multislice): https://docs.cloud.google.com/tpu/docs/system-architecture-tpu-vm ↗
- AMD MI300 node-level architecture: https://instinct.docs.amd.com/develop/gpu-arch/mi300.html ↗
- NVIDIA H100 (NVLink): https://www.nvidia.com/en-us/data-center/h100/ ↗
