RAG (retrieval-augmented generation)
**RAG** means: before (or while) generating an answer, the system **retrieves** relevant documents and feeds them into the prompt so the model can ground replies in those sources.
What it is
RAG means: before (or while) generating an answer, the system retrieves relevant documents and feeds them into the prompt so the model can ground replies in those sources.
Why it matters
Pure LLMs freeze knowledge at training time and hallucinate. RAG helps with private docs, fresh facts, and citations—when retrieval is good.
How it works (plain)
- Split your knowledge base into chunks
- Index them (often with embeddings)
- On each question, fetch top chunks
- Generate an answer using question + chunks
- Ideally cite which chunks were used—and let humans click through
Everyday example
A company help bot that searches internal policies before answering “How many PTO days do new hires get?”
Try it
Ask a chatbot a question about a PDF you paste vs one you do not. Notice how grounding changes behavior—and still verify.
Myths
- ⚠️ Myth: RAG eliminates hallucinations.
- ✓ Reality: Bad chunks or misread chunks still produce fluent errors.
- ⚠️ Myth: More retrieved text is always better.
- ✓ Reality: Noise and prompt injection risk rise (Course 29).
Sources
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP (2020): https://arxiv.org/abs/2005.11401 ↗
- ANN live: https://www.ainerdnetwork.com/learn/rag-retrieval-augmented-generation ↗
- Hugging Face LLM course (RAG sections as applicable): https://huggingface.co/learn/llm-course/ ↗
