Self-supervised speech representations (wav2vec 2.0)
**Self-supervised learning (SSL) for speech** trains a model on raw audio **without** human transcripts first, then fine-tunes on a smaller labeled set for ASR (or other tasks). **wav2vec 2.0** (Baevski et al., 2020) is a landmark framew...
What it is
Self-supervised learning (SSL) for speech trains a model on raw audio without human transcripts first, then fine-tunes on a smaller labeled set for ASR (or other tasks). wav2vec 2.0 (Baevski et al., 2020) is a landmark framework: mask latent speech features, learn quantized units, and solve a contrastive task—then fine-tune with CTC.
<!-- IMAGE: unlabeled audio → pretrained encoder → fine-tune with transcripts -->
Visual Spec & Architecture Diagram
wav2vec 2.0-style SSL diagram: raw waveform → CNN feature encoder → latent speech representations (vector sequence) → Transformer context network. Masked spans shown as gray blocks on latents. Contrastive task callout: 'identify true latent among distractors (quantized targets)'. Labels: 'self-supervised pretrain' then arrow to 'fine-tune ASR with little labeled data'.
Visual Spec & Architecture Diagram
Two scales: left mountain of unlabeled audio hours vs tiny labeled ASR hours; right arrow 'pretrain on unlabeled → adapt with labeled'. Caption: 'Why SSL matters for speech'.
Why it matters
Labeled speech is expensive. Most of the world’s languages lack thousands of transcribed hours. wav2vec 2.0 showed strong ASR with far less labeled data after pretraining on unlabeled audio—including reported ultra-low-resource settings in the paper (see Technical for numbers from the paper, not marketing).
ESPnet course materials include assignments on using self-supervised speech representations inside ASR training—SSL is now standard toolkit practice, not only a paper idea.
How it works (plain)
- Feed raw waveform into a convolutional encoder → latent speech vectors.
- Mask spans of those latents (like masked language modeling, but for audio).
- A Transformer builds context from the (partially masked) sequence.
- The model learns to pick the correct quantized latent among distractors (contrastive learning).
- Fine-tune on transcribed speech (often with CTC) for recognition.
Everyday example
You have many hours of call-center audio but only a little carefully transcribed. Pretrain (or reuse a public SSL checkpoint) on the unlabeled pile, then fine-tune on the small transcript set—instead of demanding a huge labeled corpus up front.
Try it
Read the wav2vec 2.0 abstract. Write one sentence: what is learned from unlabeled audio vs what still needs labels. Skim an ESPnet SSL assignment notebook title list to see how courses operationalize it.
Myths
- ⚠️ Myth: SSL means “zero labeled data forever.”
- ✓ Reality: Fine-tuning still uses labels; the win is *fewer* labels.
- ⚠️ Myth: Any unlabeled pile is fine.
- ✓ Reality: Domain mismatch (phone vs studio) still hurts; filter junk carefully.
- ⚠️ Myth: SSL replaces evaluation.
- ✓ Reality: You still measure WER on your target conditions.
Sources
- wav2vec 2.0 paper: https://arxiv.org/abs/2006.11477 ↗
- ESPnet notebooks (SSL assignment listed): https://espnet.github.io/espnet/notebook/ ↗
- Stanford CS224S (neural ASR homework culture): https://web.stanford.edu/class/cs224s/index.html ↗
