COURSE 06L1100% FREE
Verified 2026-08-10

MDPs and policies

A **Markov Decision Process (MDP)** is a formal story for sequential decisions: you are in a **state**, pick an **action**, get a **reward**, and land in a new state—often with randomness. A **policy** is the strategy mapping states to a...

What it is

A Markov Decision Process (MDP) is a formal story for sequential decisions: you are in a state, pick an action, get a reward, and land in a new state—often with randomness. A policy is the strategy mapping states to actions.

Why it matters

MDPs are the bridge from classical planning to reinforcement learning (Course 14). Agent loops in Course 10 are informal cousins.

How it works (plain)

  • State: where things stand
  • Action: what you can do
  • Transition: what might happen next
  • Reward: immediate feedback
  • Policy: “when I see X, do Y”

“Markov” means the future depends on the present state (as modeled), not the full ancient history—modeling choice matters.

Everyday example

A thermostat: state = temperature band; actions = heat/cool/off; reward = comfort minus energy cost.

Try it

Describe one daily routine as states/actions/rewards in five lines.

Myths

⚠️ Myth: If it is an MDP, rewards are objective truth.
✓ Reality: Reward design encodes goals—and can be gamed.
⚠️ Myth: Optimal policies are always computable cheaply.
✓ Reality: Large state spaces need approximation (RL, search).

Sources