COURSE 15L1100% FREE
Verified 2026-08-14

BART and seq2seq pretraining

**BART** is a **denoising autoencoder** for pretraining sequence-to-sequence Transformers: corrupt text with a noising function, then train the model to reconstruct the original. Lewis et al. (ACL 2020) describe it as generalizing ideas ...

What it is

BART is a denoising autoencoder for pretraining sequence-to-sequence Transformers: corrupt text with a noising function, then train the model to reconstruct the original. Lewis et al. (ACL 2020) describe it as generalizing ideas from BERT (bidirectional encoder) and GPT (left-to-right decoder) in one NMT-style architecture.

<!-- IMAGE: corrupted text → encoder-decoder → reconstructed text -->

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

BART denoising pretrain: clean text → noising operators (mask, delete, permute, sentence shuffle) → corrupted input → seq2seq model reconstructs original. Arrow to fine-tune for summarization/MT. Title: 'Denoise to learn seq2seq'.

Educational Focus: Makes BART's pretraining objective tangible vs BERT's MLM.

Why it matters

If BERT-class models shine at understanding, BART-class models shine when you need generation: summarization, dialogue, some QA formulations, and translation transfer. Choosing encoder-only vs seq2seq vs decoder-only is a core Course 15 skill.

How it works (plain)

  1. Start with clean text.
  2. Noise it (mask spans, shuffle sentences, etc.).
  3. Encode the noisy input; decode the clean original.
  4. Fine-tune on a downstream generation or understanding task.

Authors find strong results from combining sentence shuffling with a span in-filling scheme (replace spans with a single mask token).

Everyday example

Fine-tune a seq2seq model to summarize internal incident reports into a fixed template—more natural than bolting a decoder onto a pure encoder after the fact.

Try it

Skim the BART abstract. List two tasks you’d pick BART-like seq2seq for, and two you’d keep as BERT-like classification.

Myths

⚠️ Myth: Seq2seq pretraining is only for summarization.
✓ Reality: It also transfers to MT and comprehension tasks (paper evaluates both).
⚠️ Myth: Any noise function is equally good.
✓ Reality: The paper compares noising strategies; span in-filling + shuffling won for them.
⚠️ Myth: BART replaces evaluation of faithfulness.
✓ Reality: Generators still hallucinate—see summarization evaluation.

Sources