Reinforcement learning (preview)
**Reinforcement learning (RL)** learns by **trying actions** and receiving **rewards**—not by reading a full answer key for every state. It powers game agents, some robotics, and parts of how chat models are tuned with preference/reward ...
What it is
Reinforcement learning (RL) learns by trying actions and receiving rewards—not by reading a full answer key for every state.
It powers game agents, some robotics, and parts of how chat models are tuned with preference/reward signals (details in Course 14 and Course 19).
Why it matters
When the “right answer” is a sequence of decisions (play a game, control a robot, choose tool calls), RL is a natural framing.
How it works (plain)
Agent sees a situation → picks an action → gets reward (or pain) → updates policy. Explore enough to learn; exploit what works.
Everyday example
A game AI improving by playing itself. A thermostat-like controller learning to reduce energy while keeping comfort (in research/industry prototypes).
Try it
Write a reward for “clean the kitchen”: what gets +1, what gets −1? Notice how easy it is to specify the wrong goal (reward hacking foreshadowing).
Myths
- ⚠️ Myth: All modern AI is RL.
- ✓ Reality: Most production ML is supervised/self-supervised; RL is powerful but narrower in deployment count.
- ⚠️ Myth: RLHF means the chatbot “feels reward.”
- ✓ Reality: It is an optimization process using preference data—not emotions.
Sources
- Sutton & Barto, *Reinforcement Learning: An Introduction* (free draft/book site widely used in courses): http://incompleteideas.net/book/the-book-2nd.html ↗
- ANN live: https://www.ainerdnetwork.com/learn/reinforcement-learning ↗
- Course 14 for depth
