RNNs and sequence models
**Recurrent neural networks (RNNs)** process sequences step by step, passing a hidden state forward in time. **LSTMs/GRUs** are RNN variants designed to keep information longer.
What it is
Recurrent neural networks (RNNs) process sequences step by step, passing a hidden state forward in time. LSTMs/GRUs are RNN variants designed to keep information longer.
Why it matters
Before transformers dominated NLP, RNNs powered translation and speech demos. They still teach sequence thinking—and why long-range memory was hard.
How it works (plain)
At each time step: read the next token/frame, update memory, optionally emit an output. Gradients must flow backward through time, which can shrink or explode.
Everyday example
Reading a sentence word by word while keeping a running sense of the subject—until the sentence gets too long and you forget the start.
Try it
Write a 20-word sentence with a pronoun at the end referring to something at the start. That’s a long-range dependency RNNs struggled with.
Myths
- ⚠️ Myth: RNNs are obsolete and useless to learn.
- ✓ Reality: They explain history, streaming settings, and inductive biases; transformers won many benchmarks, not every niche.
- ⚠️ Myth: LSTM “solves” infinite memory.
- ✓ Reality: It helps; it does not grant perfect recall of arbitrary length.
Sources
- CS231n / classic RNN notes (verify current URL when citing)
- Course 07 transformers (why attention scaled better)
- deep learning book sequence chapters: https://www.deeplearningbook.org/ ↗
