MDP, value, and policy foundations
A **Markov decision process (MDP)** is the standard math story for sequential decisions: you are in a state, you pick an action, the world moves, you get a reward. CS234 opens with tabular MDP planning, then policy evaluation, then learn...
What it is
A Markov decision process (MDP) is the standard math story for sequential decisions: you are in a state, you pick an action, the world moves, you get a reward. CS234 opens with tabular MDP planning, then policy evaluation, then learning methods. Spinning Up’s “Key Concepts in RL” matches this vocabulary.
Visual Spec & Architecture Diagram
Classic MDP loop: Agent box with Policy π; arrow Action a_t to Environment; arrow Reward r_t and State s_{t+1} back to Agent. Labels S, A, P, R, γ on a side legend. Title: 'Markov Decision Process loop'.
Why it matters
Almost every classical RL algorithm is “find a good policy in an MDP” (or a close cousin: POMDP, contextual bandit). If MDP words are fuzzy, later deep RL and game AI lectures feel like magic.
How it works (plain)
- State: what you need to know to act.
- Action: what you can do.
- Reward: a number saying how good the last step was.
- Policy: the rule for picking actions.
- Value: how good it is to be here (or to take an action here) if you keep following a policy.
“Markov” means the future depends on the present state, not the full history—as modeled. Bad state design breaks the assumption in practice.
Everyday example
A board game position is a state; a legal move is an action; winning gives reward; a strategy book is a policy.
Try it
Pick a maze. List five states, two actions, and a reward rule. Say whether your “state” forgets anything important (keys held, etc.).
Myths
- ⚠️ Myth: Real life is always a perfect MDP.
- ✓ Reality: Partial observability and misspecified states are common—robotics and chat both stretch the model.
- ⚠️ Myth: Value functions are only for games.
- ✓ Reality: Values appear in control, operations research, and as critics inside deep RL.
Sources
- CS234 modules (tabular MDP planning; policy evaluation; SB ch. pointers): https://web.stanford.edu/class/cs234/modules.html ↗
- CS234 course description: https://web.stanford.edu/class/cs234/index.html ↗
- Spinning Up Key Concepts: https://spinningup.openai.com/ ↗
- CS285 RL Basics lecture: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
