Reinforcement Learning
Learning through interaction and reward: Markov Decision Processes (MDPs), Q-learning, policy gradients, PPO, and Reinforcement Learning from Human Feedback (RLHF).
Course Syllabus & Units
Methods
3 lessonsActor-critic overview
**Actor-critic** methods learn a policy (the **actor**) and a value helper (the **critic**) that reduces the noise of pure “multiply by total return” policy gradients. Berkeley CS285 pairs “Policy Gradients” with “Actor Critic” in the sa...
Policy gradients overview
**Policy gradient** methods learn a policy directly—nudging action probabilities in directions that raise expected reward—rather than only filling a Q-table. CS234 and CS285 both devote substantial lecture time to policy search / policy ...
Q-learning intuition
**Q-learning** learns a score for “take this action in this state”—how good that choice is, assuming you’ll act well afterward. CS234’s modules list Q-learning after tabular MDP planning and policy evaluation; Berkeley CS285 covers value...
