Edge vs cloud inference
Choosing whether models run **on-device** or in the **cloud**—latency, privacy, cost, capability, and energy tradeoffs. MLPerf Inference tracks both **edge** and **datacenter** categories/scenarios (see Inference docs model tables).
What it is
Choosing whether models run on-device or in the cloud—latency, privacy, cost, capability, and energy tradeoffs. MLPerf Inference tracks both edge and datacenter categories/scenarios (see Inference docs model tables).
Visual Spec & Architecture Diagram
Edge vs cloud inference: phone/IoT icon vs data-center; axes latency, privacy, cost, model size, connectivity. Decision arrows.
Why it matters
Not every feature should send raw audio/images to a vendor. Cloud inference aggregates into data-center electricity (IEA/LBNL); edge shifts energy to devices and can reduce round-trips.
How it works (plain)
Edge: private/fast/offline, smaller/quantized models. Cloud: bigger models/tools, data leaves device. Hybrids common (wake word local, heavy lift cloud). Measure quality on *your* tasks after quantization.
Everyday example
Offline maps on your phone vs streaming turn-by-turn from a server.
Try it
Pick one feature; argue edge vs cloud with privacy + latency + cost + energy boundary.
Myths
- ⚠️ Myth: Edge always equals private forever.
- ✓ Reality: Telemetry and updates can still exfiltrate—read policies.
- ⚠️ Myth: Cloud is always greener via “big efficient data centers.”
- ✓ Reality: Depends on grid mix, utilization, and whether idle capacity is counted—cite methods.
Sources
- MLPerf Inference: https://docs.mlcommons.org/inference/index_gh/ ↗
- IEA — Energy and AI: https://www.iea.org/reports/energy-and-ai/ ↗
- Course 24 quantization; Course 19 privacy
