COURSE 07L1100% FREE
Verified 2026-08-10

Mixture of experts (overview)

A **mixture of experts (MoE)** model has many specialist feed-forward “experts,” but only activates a few for each token. A **router** chooses which experts run.

What it is

A mixture of experts (MoE) model has many specialist feed-forward “experts,” but only activates a few for each token. A router chooses which experts run.

Why it matters

MoE is one reason some large models advertise huge parameter counts while keeping per-token compute lower than a dense model of the same size. It affects cost, latency, and failure modes (routing collapse).

How it works (plain)

Think of a hospital: not every specialist sees every patient. A triage step sends you to a subset of departments. MoE does analogous routing inside the network.

Everyday example

A huge toolbox where each screw only pulls out two wrenches—not the entire chest.

Try it

When a vendor says “trillion parameters,” ask whether the model is dense or MoE and what that means for your bill.

Myths

⚠️ Myth: More parameters always means proportionally more compute per token.
✓ Reality: Sparse MoE breaks that intuition—read the architecture notes.
⚠️ Myth: MoE always beats dense models.
✓ Reality: Training stability, routing, and hardware matter; compare on your task.

Sources

  • Provider architecture blogs for MoE models you use (cite specific posts)
  • Course 07 transformers; pretraining-and-finetuning
  • arXiv surveys on MoE (verify before quoting numbers)