Rerankers
A **reranker** takes a first-pass retrieval list and reorders it with a stronger (often slower) model so the generator sees better chunks.
What it is
A reranker takes a first-pass retrieval list and reorders it with a stronger (often slower) model so the generator sees better chunks.
Why it matters
Cheap retrieval recalls candidates; rerankers boost precision at the top—often the best RAG quality upgrade per engineering hour.
How it works (plain)
Retrieve top 50–100 → score each (query, chunk) pair → keep top k → generate with citations.
Everyday example
Search finds many pages; you skim and put the best three on top before writing a summary.
Try it
For one corpus, compare answer quality with and without a human “rerank” of chunks.
Myths
- ⚠️ Myth: Bigger embedding models remove the need to rerank.
- ✓ Reality: Cross-encoders still often win for top-k precision.
- ⚠️ Myth: Rerankers are always too slow.
- ✓ Reality: Rerank 50 short chunks is often fine in interactive apps.
Sources
- Course 09 hybrid search; groundedness
- Provider rerank API docs (cite specifically)
