Audio and speech models
A **family map** of models that handle speech and broader audio: recognition (ASR), synthesis (TTS), enhancement/separation, spoken language understanding (SLU), speech translation, music/audio generation, and non-speech event detection....
What it is
A family map of models that handle speech and broader audio: recognition (ASR), synthesis (TTS), enhancement/separation, spoken language understanding (SLU), speech translation, music/audio generation, and non-speech event detection.
<!-- IMAGE: family tree — ASR / TTS / SE / SLU / ST / events -->
Visual Spec & Architecture Diagram
Taxonomy map of speech model families as islands connected by arrows: Classical (GMM-HMM), Hybrid DNN-HMM, End-to-end CTC, Attention seq2seq, Self-supervised encoders (wav2vec/HuBERT), Multitask ASR (Whisper-style), Generative audio (TTS/codecs). Each island has 1-line purpose. Center title: 'Speech model family map'.
Why it matters
Voice UIs, accessibility, media tools, and industrial monitoring all pick *different* members of this family. Mixing them up (e.g., treating TTS as “speaker ID”) causes bad product and safety decisions. Course 29 covers voice-clone risk literacy—not cloning how-to.
How it works (plain)
Most pipelines share a pattern:
Audio in → features or a learned encoder → task head (classify, transcribe, translate) or generative decoder (speak, enhance, generate music).
Speech models often specialize on human voice. General audio models handle wider sound taxonomies (alarms, machines, animals). Modern toolkits (ESPnet notebooks, SpeechBrain, Kaldi) ship recipes for several of these tasks side by side.
Stanford CS224S ties the family to dialog and conversational systems: recognition and synthesis are pieces of assistants, not the whole product.
Everyday example
| Need | Family member |
|---|---|
| Podcast captions | ASR |
| Navigation voice | TTS |
| “Is that a siren?” | Audio event detection |
| Cleaner call audio | Speech enhancement |
| Meeting “who spoke when” | Diarization |
| Caption in another language | Speech translation |
Try it
Pick one real need from the table. Write (1) the input, (2) the desired output format, (3) what “good enough” means in one sentence. If you cannot state (3), you are not ready to pick a model.
Myths
- ⚠️ Myth: One “audio foundation model” replaces every speech product.
- ✓ Reality: Multitask models help, but latency, languages, and eval still specialize.
- ⚠️ Myth: Natural-sounding TTS proves the speaker is human.
- ✓ Reality: Synthetic speech can be highly realistic—verify high-stakes calls out-of-band (Course 29).
- ⚠️ Myth: Audio AI needs no privacy care.
- ✓ Reality: Voice is sensitive in many policies; always-on mics raise consent stakes.
Sources
- Stanford CS224S: https://web.stanford.edu/class/cs224s/index.html ↗
- ESPnet notebook index (ASR, SE, SLU, TTS, ST demos): https://espnet.github.io/espnet/notebook/ ↗
- ANN legacy page (migration pointer): https://www.ainerdnetwork.com/learn/audio-and-speech-models ↗
- Course 29 scams / voice-clone risk literacy (no how-to)
