Game RL: AlphaGo and MuZero
**AlphaGo** combined deep neural networks with search to master Go, defeating Fan Hui 5–0 (2015) and Lee Sedol 4–1 (2016). **MuZero** (DeepMind, Nature paper; blog Dec 2020) learned to master Go, chess, shogi, **and** Atari *without bein...
What it is
AlphaGo combined deep neural networks with search to master Go, defeating Fan Hui 5–0 (2015) and Lee Sedol 4–1 (2016). MuZero (DeepMind, Nature paper; blog Dec 2020) learned to master Go, chess, shogi, and Atari *without being told the rules*, by learning a model good enough for planning. CS234 includes lectures on Monte Carlo Tree Search and conquering Go.
Visual Spec & Architecture Diagram
High-level AlphaGo/MuZero literacy: self-play loop, tree search (MCTS) cloud, value/policy nets; MuZero branch 'learn model of environment dynamics' without game-cheat details. Banner: 'research systems overview—not a bot-building guide'.
Why it matters
These systems are flagship examples of classical/game RL: clear rules or simulators, self-play, search, and dense competitive rewards (win/loss). They are *not* the same setting as RLHF for chatbots—even when both use “reinforcement learning” in the name.
How it works (plain)
AlphaGo (DeepMind research page): a policy network picks moves; a value network predicts who will win; search looks ahead. It first learned from expert games, then improved by playing versions of itself (reinforcement learning).
MuZero: instead of needing a perfect rules engine for planning, it learns three planning-critical predictions—value, policy, and reward—and combines them with AlphaZero-style lookahead tree search. On Atari it set new RL state-of-the-art (per DeepMind blog); on board games it matched AlphaZero-level play.
Everyday example
AlphaGo is like a student who studies master games, then spars endlessly. MuZero is like a student who also builds a compact mental model of “what matters for planning” without memorizing every physical detail of the board room.
Try it
List three ingredients AlphaGo used (networks + search + self-play). Then write one sentence on what MuZero adds for domains where rules aren’t handed to the agent.
Myths
- ⚠️ Myth: AlphaGo “just brute-forced” Go.
- ✓ Reality: DeepMind emphasizes ~\(10^{170}\) positions—far beyond naive enumeration; learning + search were essential.
- ⚠️ Myth: MuZero models every pixel detail of the world.
- ✓ Reality: The blog stresses modeling aspects important for decisions (value, policy, reward), not full environment reconstruction.
Sources
- AlphaGo: https://deepmind.google/research/alphago/ ↗
- MuZero blog: https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules/ ↗
- CS234 MCTS / Go materials: https://web.stanford.edu/class/cs234/modules.html ↗
