Chunking and embeddings for RAG
**Chunking** splits documents into pieces a retriever can fetch. **Embeddings** turn those pieces into vectors so “similar meaning” search can work.
What it is
Chunking splits documents into pieces a retriever can fetch. Embeddings turn those pieces into vectors so “similar meaning” search can work.
Why it matters
RAG quality often dies in chunking: too big → noisy context; too small → missing glue; bad overlaps → duplicated or broken ideas.
How it works (plain)
- Split by headings, paragraphs, or token windows with overlap
- Embed each chunk
- At query time, embed the question and fetch nearest chunks
- Stuff chunks into the prompt with citations
Always show sources to humans for important answers.
Everyday example
A binder with sticky tabs: good tabs find the right section; one giant untitled pile does not.
Try it
Take one PDF page. Propose two chunking schemes and when each fails.
Myths
- ⚠️ Myth: Perfect embeddings remove the need for citations.
- ✓ Reality: Retrieval still errs; users need links.
- ⚠️ Myth: Smaller chunks are always better.
- ✓ Reality: Tiny chunks lose context; tune for the corpus.
Sources
- Course 09 rag-retrieval-augmented-generation
- OpenAI Academy / provider RAG guides (cite the one you use)
- ANN Learn RAG page
