COURSE 14L2100% FREE
Verified 2026-08-14

Actor-critic overview

**Actor-critic** methods learn a policy (the **actor**) and a value helper (the **critic**) that reduces the noise of pure “multiply by total return” policy gradients. Berkeley CS285 pairs “Policy Gradients” with “Actor Critic” in the sa...

What it is

Actor-critic methods learn a policy (the actor) and a value helper (the critic) that reduces the noise of pure “multiply by total return” policy gradients. Berkeley CS285 pairs “Policy Gradients” with “Actor Critic” in the same week; Spinning Up’s PPO is an on-policy actor-critic-style algorithm in practice.

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Actor-Critic twin networks: Actor π_θ proposes action; Critic V_w or Q_w scores; Advantage A = return − baseline arrow into actor update; TD error into critic update. Shared environment loop.

Educational Focus: Actor-critic architecture is easier as a picture than prose.

Why it matters

Many modern RL algorithms are actor-critic cousins. The pattern also appears when language-model policies are optimized against a learned reward model (RLHF)—though the *environment* is text generation, not a game board.

How it works (plain)

Actor proposes actions. Critic scores how surprising or advantageous the outcome was relative to what was expected. Actor updates using that cleaner signal. Bad rewards still produce bad behavior.

Everyday example

A player (actor) and a coach (critic): the coach doesn’t take the shots, but helps the player know which decisions were better than average.

Try it

In one sentence each, define actor and critic for a maze robot. Then note one way a bad critic could mislead the actor.

Myths

⚠️ Myth: “Critic” means a language model judging text.
✓ Reality: In classical RL, the critic is a value approximator \(V\) or \(Q\). LLM-as-judge is a different product pattern (related only loosely).
⚠️ Myth: Adding a critic removes reward hacking.
✓ Reality: Critics estimate returns under the *given* reward; they don’t invent the true human intent.

Sources