Actor-critic overview
**Actor-critic** methods learn a policy (the **actor**) and a value helper (the **critic**) that reduces the noise of pure “multiply by total return” policy gradients. Berkeley CS285 pairs “Policy Gradients” with “Actor Critic” in the sa...
What it is
Actor-critic methods learn a policy (the actor) and a value helper (the critic) that reduces the noise of pure “multiply by total return” policy gradients. Berkeley CS285 pairs “Policy Gradients” with “Actor Critic” in the same week; Spinning Up’s PPO is an on-policy actor-critic-style algorithm in practice.
Visual Spec & Architecture Diagram
Actor-Critic twin networks: Actor π_θ proposes action; Critic V_w or Q_w scores; Advantage A = return − baseline arrow into actor update; TD error into critic update. Shared environment loop.
Why it matters
Many modern RL algorithms are actor-critic cousins. The pattern also appears when language-model policies are optimized against a learned reward model (RLHF)—though the *environment* is text generation, not a game board.
How it works (plain)
Actor proposes actions. Critic scores how surprising or advantageous the outcome was relative to what was expected. Actor updates using that cleaner signal. Bad rewards still produce bad behavior.
Everyday example
A player (actor) and a coach (critic): the coach doesn’t take the shots, but helps the player know which decisions were better than average.
Try it
In one sentence each, define actor and critic for a maze robot. Then note one way a bad critic could mislead the actor.
Myths
- ⚠️ Myth: “Critic” means a language model judging text.
- ✓ Reality: In classical RL, the critic is a value approximator \(V\) or \(Q\). LLM-as-judge is a different product pattern (related only loosely).
- ⚠️ Myth: Adding a critic removes reward hacking.
- ✓ Reality: Critics estimate returns under the *given* reward; they don’t invent the true human intent.
Sources
- CS285 Lectures 5–6 (Policy Gradients; Actor Critic): https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
- Spinning Up PPO (on-policy clipped updates with advantages): https://spinningup.openai.com/en/latest/algorithms/ppo.html ↗
- CS234 policy gradient modules: https://web.stanford.edu/class/cs234/modules.html ↗
