COURSE 15L1100% FREE
Verified 2026-08-14

Summarization evaluation

**Summarization** shortens text; **summarization evaluation** judges whether the short version is faithful, complete enough, and useful—not only fluent. BART (Lewis et al., 2020) is a seq2seq pretrained model that became a strong abstrac...

What it is

Summarization shortens text; summarization evaluation judges whether the short version is faithful, complete enough, and useful—not only fluent. BART (Lewis et al., 2020) is a seq2seq pretrained model that became a strong abstractive summarization baseline; SLP3 covers generation-related topics across LLM and structure chapters.

<!-- IMAGE: long doc → summary bullets with “supported?” checkmarks -->

MEDIUM PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Split panel: extractive (highlighted source sentences) vs abstractive (new paraphrase). Metric strip: ROUGE (overlap), BERTScore (embedding), human preference. Warning icon: 'high ROUGE ≠ faithful'.

Educational Focus: Stops metric worship; visualizes extractive vs abstractive.

Why it matters

Fluent summaries can omit key risks or invent details. Leaders who “only read the summary” inherit those errors. Eval must catch hallucination and missing must-include facts.

How it works (plain)

Check four questions:

  1. Supported? Is each claim in the source?
  2. Coverage? Are must-include points present?
  3. Audience fit? Right length and jargon level?
  4. Stable? Does the system change answers under harmless rephrasing?

Prefer citation-backed summaries for serious domains (Course 09 groundedness themes).

Everyday example

Summarize a refund policy. If the summary drops the 30-day limit, it is a product bug—even if ROUGE looks fine.

Try it

Summarize a policy page. List three facts that must appear. Score a model against that list (coverage). Then hunt one unsupported sentence (faithfulness fail).

Myths

⚠️ Myth: High ROUGE means trustworthy.
✓ Reality: Overlap metrics miss hallucination.
⚠️ Myth: Abstractive always beats extractive.
✓ Reality: Extractive can be safer when inventing words is costly.
⚠️ Myth: LLM-as-judge replaces humans.
✓ Reality: Judges need spot checks (Course 18 caveats).

Sources