Mixture of experts (overview)
A **mixture of experts (MoE)** model has many specialist feed-forward “experts,” but only activates a few for each token. A **router** chooses which experts run.
What it is
A mixture of experts (MoE) model has many specialist feed-forward “experts,” but only activates a few for each token. A router chooses which experts run.
Why it matters
MoE is one reason some large models advertise huge parameter counts while keeping per-token compute lower than a dense model of the same size. It affects cost, latency, and failure modes (routing collapse).
How it works (plain)
Think of a hospital: not every specialist sees every patient. A triage step sends you to a subset of departments. MoE does analogous routing inside the network.
Everyday example
A huge toolbox where each screw only pulls out two wrenches—not the entire chest.
Try it
When a vendor says “trillion parameters,” ask whether the model is dense or MoE and what that means for your bill.
Myths
- ⚠️ Myth: More parameters always means proportionally more compute per token.
- ✓ Reality: Sparse MoE breaks that intuition—read the architecture notes.
- ⚠️ Myth: MoE always beats dense models.
- ✓ Reality: Training stability, routing, and hardware matter; compare on your task.
Sources
- Provider architecture blogs for MoE models you use (cite specific posts)
- Course 07 transformers; pretraining-and-finetuning
- arXiv surveys on MoE (verify before quoting numbers)
