COURSE 09L1100% FREE
Verified 2026-08-10

Hybrid search

**Hybrid search** mixes keyword search (exact terms, IDs, jargon) with vector/semantic search (paraphrases, meaning). A merger or re-ranker blends the lists.

What it is

Hybrid search mixes keyword search (exact terms, IDs, jargon) with vector/semantic search (paraphrases, meaning). A merger or re-ranker blends the lists.

Why it matters

Pure embeddings miss SKUs, error codes, and rare proper nouns. Pure keywords miss paraphrases. Real corpora need both.

How it works (plain)

  1. Run BM25/keyword retrieval
  2. Run dense embedding retrieval
  3. Fuse scores (e.g., reciprocal rank fusion)
  4. Optional cross-encoder re-rank
  5. Send top chunks to the generator

Everyday example

Looking up “invoice #48291” (keyword) vs “how do I dispute a late fee?” (semantic).

Try it

List 3 queries for your docs that need exact match and 3 that need paraphrase match.

Myths

⚠️ Myth: Bigger embedding models remove the need for keywords.
✓ Reality: Exact tokens still win for IDs and legalese pins.
⚠️ Myth: Hybrid always costs too much.
✓ Reality: Often cheaper than wrong answers and support tickets.

Sources

  • Course 09 RAG + chunking chapters
  • Provider search/RAG docs (cite specifically)
  • Classic BM25 IR background (Manning *Introduction to IR* — verify)