COURSE 13L1100% FREE
Verified 2026-08-14

Speech recognition overview

**Automatic speech recognition (ASR)** turns spoken audio into text. It sits next to related jobs such as wake-word detection, language ID, speaker diarization, and speech translation. Stanford’s CS224S frames spoken language technology ...

What it is

Automatic speech recognition (ASR) turns spoken audio into text. It sits next to related jobs such as wake-word detection, language ID, speaker diarization, and speech translation. Stanford’s CS224S frames spoken language technology as recognition plus synthesis, dialogue, and understanding for assistants—not ASR alone.

<!-- IMAGE: simple pipeline mic → waveform → text transcript -->

HIGH PRIORITYDIAGRAM / FLOWCHART
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Labeled left-to-right ASR pipeline diagram with 6 boxes and arrows: (1) Microphone / waveform input labeled 'raw audio (16 kHz)'; (2) Preprocessing 'VAD + noise reduce + framing'; (3) Acoustic features 'Mel spectrogram / MFCCs' with a tiny inset spectrogram sketch; (4) Acoustic model 'phones/subwords scores'; (5) Decoder 'lexicon + language model' with a beam-search fan sketch; (6) Output text 'transcript'. Caption under decoder: 'search finds best word sequence'. Side callout: 'errors compound at each stage'. Use blue for audio path, green for text output.

Educational Focus: Gives an 8th-grade map of 'sound → words' while professionals see where LM, AM, and search sit—foundation for every later speech chapter.
MEDIUM PRIORITYANNOTATED DIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Three-panel comic of the same short sentence with ASR mistakes highlighted: (A) Substitution 'their'→'there' in red; (B) Deletion with a struck word; (C) Insertion with an extra ghost word. Bottom formula strip: 'WER = (S+D+I)/N' with each letter defined. Fake transcript only.

Educational Focus: Makes Word Error Rate concrete so learners judge systems by error types, not vibes.
HIGH PRIORITYDIAGRAM / ANNOTATED CHART
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Mel spectrogram intuition: top raw waveform of a short spoken word; middle short-time Fourier spectrogram (linear Hz); bottom Mel-scaled spectrogram with warmer colors = more energy. Labels: 'time →', 'frequency (Hz) →', 'Mel bands follow human hearing (denser at low freq)'. Callout: 'ASR models often “see” Mel images, not raw samples'. Fake 'hello' utterance only.

Educational Focus: Mel features are the bridge between sound and modern ASR/TTS—learners need this sensory metaphor.

Why it matters

Captions, call centers, voice UIs, meeting notes, and accessibility tools all lean on ASR. Quality is uneven across accents, languages, noise, and jargon. When audio leaves the device, privacy and retention policy matter as much as word error rate.

How it works (plain)

  1. Capture audio (often mono, ~16 kHz for many modern systems).
  2. Turn the waveform into features (classically filterbanks / Mel spectrograms; some models learn from raw waveform).
  3. A sequence model proposes text (historically hybrid HMM+DNN systems; today often end-to-end neural models).
  4. Optional language-model or lexicon biasing rescored hypotheses.
  5. Downstream tools (search, LLM assistants, CRM) consume the transcript.

CMU’s 18-781 course description emphasizes that speech recognition mixes theoretical foundations, algorithms, and experimental practice—you need both math and careful measurement, not only a demo.

Everyday example

Dictating a text message in a moving car: road noise, proper nouns, and short words dominate errors. A quiet studio take of the same sentence often looks “solved” by comparison.

Try it

Record one sentence twice (quiet room vs noisy room). Run the same ASR tool on both. List errors by type: wrong word, missing word, added word, punctuation-only. That list is more useful than a single “accuracy %” marketing claim.

Myths

⚠️ Myth: Transcripts are courtroom-perfect.
✓ Reality: Error rates vary by domain; always verify high-stakes content with a human.
⚠️ Myth: If it works for broadcast English, it works for everyone.
✓ Reality: Language coverage, accents, and recording conditions create large gaps.
⚠️ Myth: End-to-end means “no engineering left.”
✓ Reality: Data prep, evaluation, streaming latency, and domain vocabulary still decide products.

Sources