COURSE 14L2100% FREE
Verified 2026-08-14

RLHF at a high level

**RLHF** (Reinforcement Learning from Human Feedback) fine-tunes a model using human preference data as a reward signal. InstructGPT (Ouyang et al.) describes three stages: (1) supervised fine-tuning on demonstrations, (2) train a reward...

What it is

RLHF (Reinforcement Learning from Human Feedback) fine-tunes a model using human preference data as a reward signal. InstructGPT (Ouyang et al.) describes three stages: (1) supervised fine-tuning on demonstrations, (2) train a reward model on ranked outputs, (3) optimize the policy with PPO against that reward. Earlier, Stiennon et al. applied the same pattern to summarization. Anthropic’s HH-RLHF paper applies preference modeling + RLHF to train helpful and harmless assistants.

HIGH PRIORITYFLOWCHART
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

RLHF pipeline three stages: (1) SFT on demonstrations; (2) Reward Model trained from human preference pairs (chosen > rejected); (3) RL fine-tune policy against RM with KL penalty to SFT reference. Arrows and 'human feedback' icons. Title: 'SFT → RM → RL'.

Educational Focus: Canonical RLHF storyboard for LLM literacy.

Why it matters

It explains why chat models behave differently from base LMs—without claiming RLHF “solves alignment.” CS234’s recent schedule explicitly includes offline RL / RLHF topics; Berkeley CS285 includes an LLM RL homework and lecture.

How it works (plain)

Collect comparisons → train a preference/reward model → update the policy to prefer better answers (usually with a penalty so it doesn’t drift too far) → evaluate for regressions, toxicity, and sycophancy. Related methods optimize preferences without an explicit RL loop; stacks differ by lab and year.

Everyday example

A writing coach ranks two essay drafts; you practice toward the coach’s taste—not toward a simple word-count score.

Try it

When a vendor says “aligned with RLHF,” ask: who labeled, what instructions they got, and which public evals are published.

Myths

⚠️ Myth: RLHF makes models truthful.
✓ Reality: InstructGPT improved truthfulness on some tests vs GPT-3 but still makes simple mistakes and can hallucinate.
⚠️ Myth: All chat tuning is classic three-stage RLHF.
✓ Reality: Pipelines evolve (online iterated RLHF in Anthropic; alternatives to PPO elsewhere)—read the model report.
⚠️ Myth: Helpful and harmless never conflict.
✓ Reality: Anthropic measures tension: overly “safe” answers can refuse to help; helpful-only models can be easier to red-team.

Sources