COURSE 14L2100% FREE
Verified 2026-08-14

Deep RL, policy gradients, and PPO

**Deep RL** uses neural networks as policies and/or value functions. **Proximal Policy Optimization (PPO)** is a popular on-policy algorithm that tries to take useful policy steps without updating so far that performance collapses. OpenA...

What it is

Deep RL uses neural networks as policies and/or value functions. Proximal Policy Optimization (PPO) is a popular on-policy algorithm that tries to take useful policy steps without updating so far that performance collapses. OpenAI Spinning Up documents PPO-Clip and PPO-Penalty; InstructGPT and summarization-from-feedback both fine-tune language policies with PPO.

Why it matters

PPO shows up in both classical deep RL coursework (CS285 advanced policy gradients; Spinning Up implementations) and in LLM alignment papers. Knowing what it optimizes—and that it is still on-policy RL—prevents magical thinking about “PPO = aligned.”

How it works (plain)

Collect experience with the current policy → estimate how unexpectedly good each action was → update the policy a little, with a rule that discourages huge swings → repeat. Spinning Up’s motivation: TRPO used a complex second-order constraint; PPO aims for similar stability with simpler first-order tricks (clipping or an adaptive KL penalty).

Everyday example

Practicing free throws: take a few shots (data), adjust form slightly (update), don’t change everything at once or your muscle memory collapses (clip/KL).

Try it

Read Spinning Up’s PPO “Quick Facts”: on-policy; discrete or continuous actions. Say why on-policy matters for reusing old chat logs without care.

Myths

⚠️ Myth: PPO is only for robots and Atari.
✓ Reality: Same algorithm family is used to optimize LM policies against reward models (InstructGPT).
⚠️ Myth: Clipping guarantees safety in the real world.
✓ Reality: Clipping limits *policy change per update*, not physical or social harm.

Sources