RLHF at a high level
**RLHF** (Reinforcement Learning from Human Feedback) fine-tunes a model using human preference data as a reward signal. InstructGPT (Ouyang et al.) describes three stages: (1) supervised fine-tuning on demonstrations, (2) train a reward...
What it is
RLHF (Reinforcement Learning from Human Feedback) fine-tunes a model using human preference data as a reward signal. InstructGPT (Ouyang et al.) describes three stages: (1) supervised fine-tuning on demonstrations, (2) train a reward model on ranked outputs, (3) optimize the policy with PPO against that reward. Earlier, Stiennon et al. applied the same pattern to summarization. Anthropic’s HH-RLHF paper applies preference modeling + RLHF to train helpful and harmless assistants.
Visual Spec & Architecture Diagram
RLHF pipeline three stages: (1) SFT on demonstrations; (2) Reward Model trained from human preference pairs (chosen > rejected); (3) RL fine-tune policy against RM with KL penalty to SFT reference. Arrows and 'human feedback' icons. Title: 'SFT → RM → RL'.
Why it matters
It explains why chat models behave differently from base LMs—without claiming RLHF “solves alignment.” CS234’s recent schedule explicitly includes offline RL / RLHF topics; Berkeley CS285 includes an LLM RL homework and lecture.
How it works (plain)
Collect comparisons → train a preference/reward model → update the policy to prefer better answers (usually with a penalty so it doesn’t drift too far) → evaluate for regressions, toxicity, and sycophancy. Related methods optimize preferences without an explicit RL loop; stacks differ by lab and year.
Everyday example
A writing coach ranks two essay drafts; you practice toward the coach’s taste—not toward a simple word-count score.
Try it
When a vendor says “aligned with RLHF,” ask: who labeled, what instructions they got, and which public evals are published.
Myths
- ⚠️ Myth: RLHF makes models truthful.
- ✓ Reality: InstructGPT improved truthfulness on some tests vs GPT-3 but still makes simple mistakes and can hallucinate.
- ⚠️ Myth: All chat tuning is classic three-stage RLHF.
- ✓ Reality: Pipelines evolve (online iterated RLHF in Anthropic; alternatives to PPO elsewhere)—read the model report.
- ⚠️ Myth: Helpful and harmless never conflict.
- ✓ Reality: Anthropic measures tension: overly “safe” answers can refuse to help; helpful-only models can be easier to red-team.
Sources
- InstructGPT: https://arxiv.org/abs/2203.02155 ↗
- Summarize from human feedback: https://arxiv.org/abs/2009.01325 ↗
- Anthropic HH-RLHF: https://arxiv.org/abs/2204.05862 ↗
- CS234 schedule (RLHF lecture slot): https://web.stanford.edu/class/cs234/index.html ↗
- CS285 LLM RL: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
