Game RL vs LLM RLHF (contrast)
A deliberate contrast between **classical/game reinforcement learning** (MDPs, self-play, MCTS, AlphaGo/MuZero) and **RL from human feedback for language models** (InstructGPT, summarization-from-feedback, Anthropic HH-RLHF). MASTER’s re...
What it is
A deliberate contrast between classical/game reinforcement learning (MDPs, self-play, MCTS, AlphaGo/MuZero) and RL from human feedback for language models (InstructGPT, summarization-from-feedback, Anthropic HH-RLHF). MASTER’s research guide flags this contrast as the point of Topic 4.
Visual Spec & Architecture Diagram
Contrast table-as-diagram: Game RL (clear rules, reward from win/score, self-play) vs LLM RLHF (preference rewards, human raters, KL to base model, open-ended language). Middle: both use policy optimization but different reward sources.
Why it matters
People hear “RL” and mash the stories together. Game RL usually has crisp win/loss rewards and (often) known rules or simulators. LLM RLHF optimizes a *learned preference reward* over text, with different failure modes (sycophancy, over-refusal, annotator bias).
How it works (plain)
| Lens | Game / classical RL | LLM RLHF |
|---|---|---|
| Typical goal | Win / score / control | Helpful, honest, harmless-ish text |
| Reward | Environment score or win/loss | Reward model from human rankings |
| Data | Self-play, simulators, demos | Demonstrations + comparisons |
| Search | Often MCTS / planning | Usually next-token policy optimization (e.g. PPO), not board search |
| Famous cases | AlphaGo, MuZero | InstructGPT, Stiennon et al., HH-RLHF |
Shared math tools (policies, advantages, PPO) can appear in both—problem setting still differs.
Everyday example
Coaching a chess player with a rating system vs coaching a writing assistant with pairwise editor preferences—both “feedback,” different sports.
Try it
Write two bullets: one risk unique to game RL (e.g. exploiting simulator bugs), one unique to RLHF (e.g. reward-model overoptimization). Circle which setting a product blog is actually describing.
Myths
- ⚠️ Myth: MuZero’s “without rules” means chatbots don’t need preference data.
- ✓ Reality: MuZero still learns from environment interaction and rewards; LLMs need human (or AI) preference signals for RLHF-style steering.
- ⚠️ Myth: If it uses PPO, it’s “doing AlphaGo.”
- ✓ Reality: PPO is an optimizer; AlphaGo’s public story centers on policy/value nets + search + self-play in Go.
Sources
- AlphaGo: https://deepmind.google/research/alphago/ ↗
- MuZero: https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules/ ↗
- InstructGPT: https://arxiv.org/abs/2203.02155 ↗
- Summarize from HF: https://arxiv.org/abs/2009.01325 ↗
- HH-RLHF: https://arxiv.org/abs/2204.05862 ↗
- CS234 (MCTS + RLHF both in syllabus): https://web.stanford.edu/class/cs234/index.html ↗
- CS285 (game-style deep RL + LLM RL homework): https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
