COURSE 13L1100% FREE
Verified 2026-08-14

Self-supervised speech representations (wav2vec 2.0)

**Self-supervised learning (SSL) for speech** trains a model on raw audio **without** human transcripts first, then fine-tunes on a smaller labeled set for ASR (or other tasks). **wav2vec 2.0** (Baevski et al., 2020) is a landmark framew...

What it is

Self-supervised learning (SSL) for speech trains a model on raw audio without human transcripts first, then fine-tunes on a smaller labeled set for ASR (or other tasks). wav2vec 2.0 (Baevski et al., 2020) is a landmark framework: mask latent speech features, learn quantized units, and solve a contrastive task—then fine-tune with CTC.

<!-- IMAGE: unlabeled audio → pretrained encoder → fine-tune with transcripts -->

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

wav2vec 2.0-style SSL diagram: raw waveform → CNN feature encoder → latent speech representations (vector sequence) → Transformer context network. Masked spans shown as gray blocks on latents. Contrastive task callout: 'identify true latent among distractors (quantized targets)'. Labels: 'self-supervised pretrain' then arrow to 'fine-tune ASR with little labeled data'.

Educational Focus: SSL is abstract; this diagram is the chapter's learning core for General+Technical tabs.
MEDIUM PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Two scales: left mountain of unlabeled audio hours vs tiny labeled ASR hours; right arrow 'pretrain on unlabeled → adapt with labeled'. Caption: 'Why SSL matters for speech'.

Educational Focus: Motivates SSL economically and scientifically in one glance.

Why it matters

Labeled speech is expensive. Most of the world’s languages lack thousands of transcribed hours. wav2vec 2.0 showed strong ASR with far less labeled data after pretraining on unlabeled audio—including reported ultra-low-resource settings in the paper (see Technical for numbers from the paper, not marketing).

ESPnet course materials include assignments on using self-supervised speech representations inside ASR training—SSL is now standard toolkit practice, not only a paper idea.

How it works (plain)

  1. Feed raw waveform into a convolutional encoder → latent speech vectors.
  2. Mask spans of those latents (like masked language modeling, but for audio).
  3. A Transformer builds context from the (partially masked) sequence.
  4. The model learns to pick the correct quantized latent among distractors (contrastive learning).
  5. Fine-tune on transcribed speech (often with CTC) for recognition.

Everyday example

You have many hours of call-center audio but only a little carefully transcribed. Pretrain (or reuse a public SSL checkpoint) on the unlabeled pile, then fine-tune on the small transcript set—instead of demanding a huge labeled corpus up front.

Try it

Read the wav2vec 2.0 abstract. Write one sentence: what is learned from unlabeled audio vs what still needs labels. Skim an ESPnet SSL assignment notebook title list to see how courses operationalize it.

Myths

⚠️ Myth: SSL means “zero labeled data forever.”
✓ Reality: Fine-tuning still uses labels; the win is *fewer* labels.
⚠️ Myth: Any unlabeled pile is fine.
✓ Reality: Domain mismatch (phone vs studio) still hurts; filter junk carefully.
⚠️ Myth: SSL replaces evaluation.
✓ Reality: You still measure WER on your target conditions.

Sources