COURSE 14L2100% FREE
Verified 2026-08-14

MDP, value, and policy foundations

A **Markov decision process (MDP)** is the standard math story for sequential decisions: you are in a state, you pick an action, the world moves, you get a reward. CS234 opens with tabular MDP planning, then policy evaluation, then learn...

What it is

A Markov decision process (MDP) is the standard math story for sequential decisions: you are in a state, you pick an action, the world moves, you get a reward. CS234 opens with tabular MDP planning, then policy evaluation, then learning methods. Spinning Up’s “Key Concepts in RL” matches this vocabulary.

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Classic MDP loop: Agent box with Policy π; arrow Action a_t to Environment; arrow Reward r_t and State s_{t+1} back to Agent. Labels S, A, P, R, γ on a side legend. Title: 'Markov Decision Process loop'.

Educational Focus: MDP loop is the universal diagram for the whole course.

Why it matters

Almost every classical RL algorithm is “find a good policy in an MDP” (or a close cousin: POMDP, contextual bandit). If MDP words are fuzzy, later deep RL and game AI lectures feel like magic.

How it works (plain)

  • State: what you need to know to act.
  • Action: what you can do.
  • Reward: a number saying how good the last step was.
  • Policy: the rule for picking actions.
  • Value: how good it is to be here (or to take an action here) if you keep following a policy.

“Markov” means the future depends on the present state, not the full history—as modeled. Bad state design breaks the assumption in practice.

Everyday example

A board game position is a state; a legal move is an action; winning gives reward; a strategy book is a policy.

Try it

Pick a maze. List five states, two actions, and a reward rule. Say whether your “state” forgets anything important (keys held, etc.).

Myths

⚠️ Myth: Real life is always a perfect MDP.
✓ Reality: Partial observability and misspecified states are common—robotics and chat both stretch the model.
⚠️ Myth: Value functions are only for games.
✓ Reality: Values appear in control, operations research, and as critics inside deep RL.

Sources