Reinforcement learning overview
**Reinforcement learning (RL)** trains an agent through trial and error: the agent takes actions, gets rewards over time, and improves a policy. Stanford CS234 describes RL as a paradigm for learning to make good decisions in tasks spann...
What it is
Reinforcement learning (RL) trains an agent through trial and error: the agent takes actions, gets rewards over time, and improves a policy. Stanford CS234 describes RL as a paradigm for learning to make good decisions in tasks spanning robotics, games, consumer modeling, and healthcare.
Visual Spec & Architecture Diagram
Three columns SL / UL / RL: data type, feedback type, goal. RL column shows agent-environment loop thumbnail.
Why it matters
RL ideas show up in game AI, robotics, recommendations, and preference-tuning stages for language models (RLHF). A wrong reward still “wins” on the score while failing the real goal.
How it works (plain)
Agent sees a state → picks an action → gets a reward and a next state → updates how it chooses actions. Exploration (try new things) vs exploitation (use what already works) never fully goes away. Courses such as Berkeley CS285 start from imitation and RL basics, then move into deep policy and value methods.
Everyday example
Learning a new board game without reading the full rules—messy at first, then sharper strategies as wins and losses teach you.
Try it
Write a one-sentence reward for a “helpful study bot.” Then write one way the bot could game that reward without really helping.
Myths
- ⚠️ Myth: RL is how all modern chatbots are trained end-to-end.
- ✓ Reality: Most LLM stacks are pretrain + supervised fine-tune heavy; RL-style preference optimization is usually one later stage (see InstructGPT / HH-RLHF).
- ⚠️ Myth: Positive reward means ethical behavior.
- ✓ Reality: Rewards encode a measured goal; they can omit safety, honesty, or fairness.
Sources
- Stanford CS234: https://web.stanford.edu/class/cs234/index.html ↗
- Berkeley CS285 Deep RL: https://rail.eecs.berkeley.edu/deeprlcourse/index.html ↗
- OpenAI Spinning Up: https://spinningup.openai.com/ ↗
- Sutton & Barto (classic text often paired with CS234 modules): https://www.incompleteideas.net/book/the-book-2nd.html ↗
