Summarization evaluation
**Summarization** shortens text; **summarization evaluation** judges whether the short version is faithful, complete enough, and useful—not only fluent. BART (Lewis et al., 2020) is a seq2seq pretrained model that became a strong abstrac...
What it is
Summarization shortens text; summarization evaluation judges whether the short version is faithful, complete enough, and useful—not only fluent. BART (Lewis et al., 2020) is a seq2seq pretrained model that became a strong abstractive summarization baseline; SLP3 covers generation-related topics across LLM and structure chapters.
<!-- IMAGE: long doc → summary bullets with “supported?” checkmarks -->
Visual Spec & Architecture Diagram
Split panel: extractive (highlighted source sentences) vs abstractive (new paraphrase). Metric strip: ROUGE (overlap), BERTScore (embedding), human preference. Warning icon: 'high ROUGE ≠ faithful'.
Why it matters
Fluent summaries can omit key risks or invent details. Leaders who “only read the summary” inherit those errors. Eval must catch hallucination and missing must-include facts.
How it works (plain)
Check four questions:
- Supported? Is each claim in the source?
- Coverage? Are must-include points present?
- Audience fit? Right length and jargon level?
- Stable? Does the system change answers under harmless rephrasing?
Prefer citation-backed summaries for serious domains (Course 09 groundedness themes).
Everyday example
Summarize a refund policy. If the summary drops the 30-day limit, it is a product bug—even if ROUGE looks fine.
Try it
Summarize a policy page. List three facts that must appear. Score a model against that list (coverage). Then hunt one unsupported sentence (faithfulness fail).
Myths
- ⚠️ Myth: High ROUGE means trustworthy.
- ✓ Reality: Overlap metrics miss hallucination.
- ⚠️ Myth: Abstractive always beats extractive.
- ✓ Reality: Extractive can be safer when inventing words is costly.
- ⚠️ Myth: LLM-as-judge replaces humans.
- ✓ Reality: Judges need spot checks (Course 18 caveats).
Sources
- BART paper (summarization SOTA claims in abstract): https://aclanthology.org/2020.acl-main.703/ ↗
- SLP3: https://web.stanford.edu/~jurafsky/slp3/ ↗
- Course 08 hallucinations; Course 09 groundedness; Course 18 eval
