Speech recognition overview
**Automatic speech recognition (ASR)** turns spoken audio into text. It sits next to related jobs such as wake-word detection, language ID, speaker diarization, and speech translation. Stanford’s CS224S frames spoken language technology ...
What it is
Automatic speech recognition (ASR) turns spoken audio into text. It sits next to related jobs such as wake-word detection, language ID, speaker diarization, and speech translation. Stanford’s CS224S frames spoken language technology as recognition plus synthesis, dialogue, and understanding for assistants—not ASR alone.
<!-- IMAGE: simple pipeline mic → waveform → text transcript -->
Visual Spec & Architecture Diagram
Labeled left-to-right ASR pipeline diagram with 6 boxes and arrows: (1) Microphone / waveform input labeled 'raw audio (16 kHz)'; (2) Preprocessing 'VAD + noise reduce + framing'; (3) Acoustic features 'Mel spectrogram / MFCCs' with a tiny inset spectrogram sketch; (4) Acoustic model 'phones/subwords scores'; (5) Decoder 'lexicon + language model' with a beam-search fan sketch; (6) Output text 'transcript'. Caption under decoder: 'search finds best word sequence'. Side callout: 'errors compound at each stage'. Use blue for audio path, green for text output.
Visual Spec & Architecture Diagram
Three-panel comic of the same short sentence with ASR mistakes highlighted: (A) Substitution 'their'→'there' in red; (B) Deletion with a struck word; (C) Insertion with an extra ghost word. Bottom formula strip: 'WER = (S+D+I)/N' with each letter defined. Fake transcript only.
Visual Spec & Architecture Diagram
Mel spectrogram intuition: top raw waveform of a short spoken word; middle short-time Fourier spectrogram (linear Hz); bottom Mel-scaled spectrogram with warmer colors = more energy. Labels: 'time →', 'frequency (Hz) →', 'Mel bands follow human hearing (denser at low freq)'. Callout: 'ASR models often “see” Mel images, not raw samples'. Fake 'hello' utterance only.
Why it matters
Captions, call centers, voice UIs, meeting notes, and accessibility tools all lean on ASR. Quality is uneven across accents, languages, noise, and jargon. When audio leaves the device, privacy and retention policy matter as much as word error rate.
How it works (plain)
- Capture audio (often mono, ~16 kHz for many modern systems).
- Turn the waveform into features (classically filterbanks / Mel spectrograms; some models learn from raw waveform).
- A sequence model proposes text (historically hybrid HMM+DNN systems; today often end-to-end neural models).
- Optional language-model or lexicon biasing rescored hypotheses.
- Downstream tools (search, LLM assistants, CRM) consume the transcript.
CMU’s 18-781 course description emphasizes that speech recognition mixes theoretical foundations, algorithms, and experimental practice—you need both math and careful measurement, not only a demo.
Everyday example
Dictating a text message in a moving car: road noise, proper nouns, and short words dominate errors. A quiet studio take of the same sentence often looks “solved” by comparison.
Try it
Record one sentence twice (quiet room vs noisy room). Run the same ASR tool on both. List errors by type: wrong word, missing word, added word, punctuation-only. That list is more useful than a single “accuracy %” marketing claim.
Myths
- ⚠️ Myth: Transcripts are courtroom-perfect.
- ✓ Reality: Error rates vary by domain; always verify high-stakes content with a human.
- ⚠️ Myth: If it works for broadcast English, it works for everyone.
- ✓ Reality: Language coverage, accents, and recording conditions create large gaps.
- ⚠️ Myth: End-to-end means “no engineering left.”
- ✓ Reality: Data prep, evaluation, streaming latency, and domain vocabulary still decide products.
Sources
- Stanford CS224S (Spoken Language Processing): https://web.stanford.edu/class/cs224s/index.html ↗
- CMU 18-781 Speech Recognition and Understanding: https://courses.ece.cmu.edu/18781 ↗
- Jurafsky & Martin, *Speech and Language Processing* (3rd ed. draft) ASR chapters: https://web.stanford.edu/~jurafsky/slp3/ ↗
- NIST AI RMF (risk framing): https://www.nist.gov/itl/ai-risk-management-framework ↗
