COURSE 14L2100% FREE
Verified 2026-08-14

Q-learning intuition

**Q-learning** learns a score for “take this action in this state”—how good that choice is, assuming you’ll act well afterward. CS234’s modules list Q-learning after tabular MDP planning and policy evaluation; Berkeley CS285 covers value...

What it is

Q-learning learns a score for “take this action in this state”—how good that choice is, assuming you’ll act well afterward. CS234’s modules list Q-learning after tabular MDP planning and policy evaluation; Berkeley CS285 covers value-based RL and Q-learning in practice.

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Q-learning update intuition: table cell Q(s,a) highlighted; arrow showing old value blended with 'r + γ max Q(s',·)'; speech bubble 'bootstrap from best next action'. Tiny 3×3 gridworld with one update step numbered.

Educational Focus: The TD update is the Technical heart—must be visual.

Why it matters

It is a classic foundation before deep RL. Once you feel a Bellman update by hand, DQN-style function approximation is easier to understand.

How it works (plain)

Try actions → see reward and next state → nudge the Q number toward “reward + discounted best next Q.” Sometimes explore randomly so you don’t get stuck on an early lucky path.

Everyday example

A restaurant rating for “order the special on Tuesday”—the score depends on today’s meal *and* how good future visits look if you keep going there.

Try it

Draw a 3-state toy grid on paper. After one transition, update one Q value with a small learning rate.

Myths

⚠️ Myth: Q-tables scale to every problem.
✓ Reality: Huge state spaces need function approximation (deep Q networks and cousins).
⚠️ Myth: Higher Q always means safer action.
✓ Reality: Q reflects the reward you defined—not ethics or physical safety.

Sources