BART and seq2seq pretraining
**BART** is a **denoising autoencoder** for pretraining sequence-to-sequence Transformers: corrupt text with a noising function, then train the model to reconstruct the original. Lewis et al. (ACL 2020) describe it as generalizing ideas ...
What it is
BART is a denoising autoencoder for pretraining sequence-to-sequence Transformers: corrupt text with a noising function, then train the model to reconstruct the original. Lewis et al. (ACL 2020) describe it as generalizing ideas from BERT (bidirectional encoder) and GPT (left-to-right decoder) in one NMT-style architecture.
<!-- IMAGE: corrupted text → encoder-decoder → reconstructed text -->
Visual Spec & Architecture Diagram
BART denoising pretrain: clean text → noising operators (mask, delete, permute, sentence shuffle) → corrupted input → seq2seq model reconstructs original. Arrow to fine-tune for summarization/MT. Title: 'Denoise to learn seq2seq'.
Why it matters
If BERT-class models shine at understanding, BART-class models shine when you need generation: summarization, dialogue, some QA formulations, and translation transfer. Choosing encoder-only vs seq2seq vs decoder-only is a core Course 15 skill.
How it works (plain)
- Start with clean text.
- Noise it (mask spans, shuffle sentences, etc.).
- Encode the noisy input; decode the clean original.
- Fine-tune on a downstream generation or understanding task.
Authors find strong results from combining sentence shuffling with a span in-filling scheme (replace spans with a single mask token).
Everyday example
Fine-tune a seq2seq model to summarize internal incident reports into a fixed template—more natural than bolting a decoder onto a pure encoder after the fact.
Try it
Skim the BART abstract. List two tasks you’d pick BART-like seq2seq for, and two you’d keep as BERT-like classification.
Myths
- ⚠️ Myth: Seq2seq pretraining is only for summarization.
- ✓ Reality: It also transfers to MT and comprehension tasks (paper evaluates both).
- ⚠️ Myth: Any noise function is equally good.
- ✓ Reality: The paper compares noising strategies; span in-filling + shuffling won for them.
- ⚠️ Myth: BART replaces evaluation of faithfulness.
- ✓ Reality: Generators still hallucinate—see summarization evaluation.
Sources
- BART paper: https://aclanthology.org/2020.acl-main.703/ ↗
- BERT (encoder contrast): https://aclanthology.org/N19-1423/ ↗
- SLP3: https://web.stanford.edu/~jurafsky/slp3/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
