Policy gradients overview
**Policy gradient** methods learn a policy directly—nudging action probabilities in directions that raise expected reward—rather than only filling a Q-table. CS234 and CS285 both devote substantial lecture time to policy search / policy ...
What it is
Policy gradient methods learn a policy directly—nudging action probabilities in directions that raise expected reward—rather than only filling a Q-table. CS234 and CS285 both devote substantial lecture time to policy search / policy gradients; Spinning Up’s “Intro to Policy Optimization” derives the simplest policy gradient and reward-to-go variants.
Visual Spec & Architecture Diagram
Policy gradient sketch: neural net policy π_θ(a|s) outputting action probs; sampled trajectory; return G; arrow 'increase prob of actions on high-return trajectories / decrease on low'. Formula callout ∇J ≈ E[G ∇log π] without requiring derivation mastery.
Why it matters
Many modern deep RL algorithms (including PPO, used in InstructGPT’s RL stage) sit in this family. Understanding the idea helps you read papers and demystify preference-tuning stacks without claiming all chat training is pure RL.
How it works (plain)
Sample actions from the current policy → see how the episode went → slightly increase probabilities of actions that led to better returns (and decrease others). Without tricks, the signal is noisy—so people use baselines, advantages, and trust-region/clipping methods.
Everyday example
A street musician adjusting how often they play each song based on tips—noisy feedback, but the mix slowly shifts.
Try it
Name policy, reward, and one failure mode (reward hacking) for a study-helper agent. Link to Course 06 MDPs if you have that chapter open.
Myths
- ⚠️ Myth: Policy gradients always beat Q-learning.
- ✓ Reality: Problem-dependent; actor-critic hybrids are common (CS285 weeks on both).
- ⚠️ Myth: “Policy gradient” means the same thing as RLHF.
- ✓ Reality: RLHF often *uses* a policy-gradient-style update (e.g. PPO) against a *learned preference reward*—a different problem setting than games.
Sources
- CS234 policy gradient / policy search materials: https://web.stanford.edu/class/cs234/modules.html ↗
- CS285 Lectures 5–6, 9–10: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
- Spinning Up Part 3: https://spinningup.openai.com/ ↗
- PPO page (practical modern PG algorithm): https://spinningup.openai.com/en/latest/algorithms/ppo.html ↗
