Deep RL, policy gradients, and PPO
**Deep RL** uses neural networks as policies and/or value functions. **Proximal Policy Optimization (PPO)** is a popular on-policy algorithm that tries to take useful policy steps without updating so far that performance collapses. OpenA...
What it is
Deep RL uses neural networks as policies and/or value functions. Proximal Policy Optimization (PPO) is a popular on-policy algorithm that tries to take useful policy steps without updating so far that performance collapses. OpenAI Spinning Up documents PPO-Clip and PPO-Penalty; InstructGPT and summarization-from-feedback both fine-tune language policies with PPO.
Why it matters
PPO shows up in both classical deep RL coursework (CS285 advanced policy gradients; Spinning Up implementations) and in LLM alignment papers. Knowing what it optimizes—and that it is still on-policy RL—prevents magical thinking about “PPO = aligned.”
How it works (plain)
Collect experience with the current policy → estimate how unexpectedly good each action was → update the policy a little, with a rule that discourages huge swings → repeat. Spinning Up’s motivation: TRPO used a complex second-order constraint; PPO aims for similar stability with simpler first-order tricks (clipping or an adaptive KL penalty).
Everyday example
Practicing free throws: take a few shots (data), adjust form slightly (update), don’t change everything at once or your muscle memory collapses (clip/KL).
Try it
Read Spinning Up’s PPO “Quick Facts”: on-policy; discrete or continuous actions. Say why on-policy matters for reusing old chat logs without care.
Myths
- ⚠️ Myth: PPO is only for robots and Atari.
- ✓ Reality: Same algorithm family is used to optimize LM policies against reward models (InstructGPT).
- ⚠️ Myth: Clipping guarantees safety in the real world.
- ✓ Reality: Clipping limits *policy change per update*, not physical or social harm.
Sources
- Spinning Up PPO: https://spinningup.openai.com/en/latest/algorithms/ppo.html ↗
- Spinning Up home: https://spinningup.openai.com/ ↗
- CS285 Deep RL course: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
- InstructGPT (PPO stage): https://arxiv.org/abs/2203.02155 ↗
