COURSE 09L1100% FREE
Verified 2026-08-10

RAG (retrieval-augmented generation)

**RAG** means: before (or while) generating an answer, the system **retrieves** relevant documents and feeds them into the prompt so the model can ground replies in those sources.

What it is

RAG means: before (or while) generating an answer, the system retrieves relevant documents and feeds them into the prompt so the model can ground replies in those sources.

Why it matters

Pure LLMs freeze knowledge at training time and hallucinate. RAG helps with private docs, fresh facts, and citations—when retrieval is good.

How it works (plain)

  1. Split your knowledge base into chunks
  2. Index them (often with embeddings)
  3. On each question, fetch top chunks
  4. Generate an answer using question + chunks
  5. Ideally cite which chunks were used—and let humans click through

Everyday example

A company help bot that searches internal policies before answering “How many PTO days do new hires get?”

Try it

Ask a chatbot a question about a PDF you paste vs one you do not. Notice how grounding changes behavior—and still verify.

Myths

⚠️ Myth: RAG eliminates hallucinations.
✓ Reality: Bad chunks or misread chunks still produce fluent errors.
⚠️ Myth: More retrieved text is always better.
✓ Reality: Noise and prompt injection risk rise (Course 29).

Sources