Q-learning intuition
**Q-learning** learns a score for “take this action in this state”—how good that choice is, assuming you’ll act well afterward. CS234’s modules list Q-learning after tabular MDP planning and policy evaluation; Berkeley CS285 covers value...
What it is
Q-learning learns a score for “take this action in this state”—how good that choice is, assuming you’ll act well afterward. CS234’s modules list Q-learning after tabular MDP planning and policy evaluation; Berkeley CS285 covers value-based RL and Q-learning in practice.
Visual Spec & Architecture Diagram
Q-learning update intuition: table cell Q(s,a) highlighted; arrow showing old value blended with 'r + γ max Q(s',·)'; speech bubble 'bootstrap from best next action'. Tiny 3×3 gridworld with one update step numbered.
Why it matters
It is a classic foundation before deep RL. Once you feel a Bellman update by hand, DQN-style function approximation is easier to understand.
How it works (plain)
Try actions → see reward and next state → nudge the Q number toward “reward + discounted best next Q.” Sometimes explore randomly so you don’t get stuck on an early lucky path.
Everyday example
A restaurant rating for “order the special on Tuesday”—the score depends on today’s meal *and* how good future visits look if you keep going there.
Try it
Draw a 3-state toy grid on paper. After one transition, update one Q value with a small learning rate.
Myths
- ⚠️ Myth: Q-tables scale to every problem.
- ✓ Reality: Huge state spaces need function approximation (deep Q networks and cousins).
- ⚠️ Myth: Higher Q always means safer action.
- ✓ Reality: Q reflects the reward you defined—not ethics or physical safety.
Sources
- CS234 lecture materials (Q-learning; Sutton & Barto ch. pointers on modules page): https://web.stanford.edu/class/cs234/modules.html ↗
- CS285 value-based RL / Q-learning in practice: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
- Sutton & Barto: https://www.incompleteideas.net/book/the-book-2nd.html ↗
