Classic information retrieval
**Information retrieval (IR)** ranks documents for a query. Classic IR uses inverted indexes and lexical scoring (TF-IDF, BM25). SLP3 includes an **Information Retrieval and Retrieval-Augmented Generation** chapter—signal that keyword se...
What it is
Information retrieval (IR) ranks documents for a query. Classic IR uses inverted indexes and lexical scoring (TF-IDF, BM25). SLP3 includes an Information Retrieval and Retrieval-Augmented Generation chapter—signal that keyword search and RAG share one lineage. CS224N’s modern NLP stack still sits on top of “find the right text first.”
<!-- IMAGE: query → inverted index → ranked hits -->
Visual Spec & Architecture Diagram
Inverted index diagram: dictionary terms 'cat','dog','ai' each pointing to posting lists of doc IDs with fake tf. Query 'cat AND ai' intersecting lists → ranked results. Side: bag-of-words vs dense retrieval note 'classic IR baseline'.
Why it matters
Keyword search still saves products when users type IDs, error codes, and exact phrases. Pure vector search misses those. Hybrid retrieval (Course 09) combines both.
How it works (plain)
- Index tokens → postings lists (which docs contain which terms).
- Score documents for a query (BM25-style).
- Return top‑k; evaluate rankings with labeled relevance or online metrics.
- Optionally expand queries or rerank with neural models.
- For RAG: retrieve first, then generate with citations.
Everyday example
Searching INV-001234 in a help center—lexical match wins; fuzzy semantic search may wander.
Try it
Write five queries that need exact match and five that need paraphrase. That split designs your hybrid system.
Myths
- ⚠️ Myth: Semantic search makes keywords obsolete.
- ✓ Reality: Hybrid wins for real corpora.
- ⚠️ Myth: IR is “solved,” only LLMs matter now.
- ✓ Reality: Bad retrieval makes good generators invent.
- ⚠️ Myth: More chunks always help RAG.
- ✓ Reality: Noise in context hurts faithfulness.
Sources
- SLP3 IR/RAG chapter: https://web.stanford.edu/~jurafsky/slp3/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
- Course 09 hybrid-search / evaluating-rag-systems
