COURSE 03L1100% FREE
Verified 2026-08-10

Reinforcement learning (preview)

**Reinforcement learning (RL)** learns by **trying actions** and receiving **rewards**—not by reading a full answer key for every state. It powers game agents, some robotics, and parts of how chat models are tuned with preference/reward ...

What it is

Reinforcement learning (RL) learns by trying actions and receiving rewards—not by reading a full answer key for every state.

It powers game agents, some robotics, and parts of how chat models are tuned with preference/reward signals (details in Course 14 and Course 19).

Why it matters

When the “right answer” is a sequence of decisions (play a game, control a robot, choose tool calls), RL is a natural framing.

How it works (plain)

Agent sees a situation → picks an action → gets reward (or pain) → updates policy. Explore enough to learn; exploit what works.

Everyday example

A game AI improving by playing itself. A thermostat-like controller learning to reduce energy while keeping comfort (in research/industry prototypes).

Try it

Write a reward for “clean the kitchen”: what gets +1, what gets −1? Notice how easy it is to specify the wrong goal (reward hacking foreshadowing).

Myths

⚠️ Myth: All modern AI is RL.
✓ Reality: Most production ML is supervised/self-supervised; RL is powerful but narrower in deployment count.
⚠️ Myth: RLHF means the chatbot “feels reward.”
✓ Reality: It is an optimization process using preference data—not emotions.

Sources