Multilingual ASR and Whisper
**Whisper** (Radford et al., 2022) is a family of encoder–decoder speech models trained with **large-scale weak supervision**: predict transcripts (and related tasks) from internet-scale audio–transcript pairs. The headline idea: scale m...
What it is
Whisper (Radford et al., 2022) is a family of encoder–decoder speech models trained with large-scale weak supervision: predict transcripts (and related tasks) from internet-scale audio–transcript pairs. The headline idea: scale multilingual, multitask supervised training so models transfer zero-shot to many benchmarks without per-dataset fine-tuning.
<!-- IMAGE: 30s audio chunk → encoder → decoder tokens (language, task, text) -->
Visual Spec & Architecture Diagram
Whisper-style multitask token diagram: audio encoder output → decoder that consumes special tokens in order: <|startoftranscript|> → <|en|> language → <|transcribe|> or <|translate|> → <|notimestamps|> → text tokens → <|endoftranscript|>. Show alternate branch for translate to English. Fake transcript 'Hello world'. Title: 'Multitask special tokens steer one model'.
Visual Spec & Architecture Diagram
Simple flowchart: audio → language ID token → ASR path in that language; wrong LID = wrong transcript comic bubble.
Why it matters
Many teams need “good enough” ASR across languages and noisy conditions without building a custom decoder per domain. Whisper popularized releasing inference code and model sizes as a shared baseline. It is not magic: transcript filtering, language coverage, and hallucination modes still matter.
How it works (plain)
- Collect lots of audio paired with transcripts on the web; filter low-quality / machine-looking captions.
- Train a sequence-to-sequence Transformer to decode text tokens from Mel spectrogram audio.
- Use special tokens to specify language, task (transcribe vs translate), timestamps, etc.
- At inference, run the model zero-shot on new audio; optionally fine-tune later for a niche domain.
Authors report training on 680,000 hours of weakly supervised audio; of that, 117,000 hours cover 96 languages beyond English, plus 125,000 hours of X→English translation data (paper-reported).
Everyday example
You paste a French interview into a tool powered by a Whisper-class model and get a usable English gist (speech translation mode) or a French transcript—without training on that interview’s domain first. Still spot-check names and numbers.
Try it
Transcribe the same clip with an English-only vs multilingual model size (tiny vs larger, if available). Compare errors on: (1) proper nouns, (2) code-switching, (3) background music. Read the paper’s caution about speaker-name hallucination tendencies during development.
Myths
- ⚠️ Myth: Zero-shot means zero errors.
- ✓ Reality: It means no fine-tune step—not perfect transcripts.
- ⚠️ Myth: Weak supervision is “dirty data, so worse than SSL.”
- ✓ Reality: Whisper argues careful filtering + scale can yield robust out-of-the-box ASR; different tradeoffs than wav2vec-style SSL.
- ⚠️ Myth: One model size fits all devices.
- ✓ Reality: Tiny→large parameter tiers trade accuracy for compute (see paper Table 1 family).
Sources
- Whisper paper: https://arxiv.org/abs/2212.04356 ↗
- Stanford CS224S (fine-tuning non-English speech homework theme): https://web.stanford.edu/class/cs224s/index.html ↗
- Related multitask speech tooling: https://espnet.github.io/espnet/notebook/ ↗
