Text-to-speech basics
**Text-to-speech (TTS)** turns written text into spoken audio. Modern neural TTS can sound natural. The same progress raises impersonation risk—this chapter teaches **how TTS systems are built and evaluated**, not how to clone a person w...
What it is
Text-to-speech (TTS) turns written text into spoken audio. Modern neural TTS can sound natural. The same progress raises impersonation risk—this chapter teaches how TTS systems are built and evaluated, not how to clone a person without consent. For scam literacy, see Course 13 unit-05 and Course 29.
<!-- IMAGE: text → linguistic features → spectrogram → waveform -->
Visual Spec & Architecture Diagram
TTS pipeline: Text → Text normalization → Linguistic features / phonemes → Acoustic model (mel spectrogram) → Vocoder (waveform) → Speaker. Show mel heatmap between acoustic model and vocoder. Labels: 'duration', 'pitch/prosody', 'vocoder reconstructs audio'. Optional dual path: concatenative vs neural TTS as faded alternatives.
Visual Spec & Architecture Diagram
UI wireframe of a TTS playground: text box, sliders for speaking rate / pitch / emotion (neutral, happy, calm), speaker dropdown with fake names 'Speaker A/B', waveform preview. Labels call out 'prosody ≠ content'.
Why it matters
Accessibility (screen readers), IVR phone trees, education, and creative tools benefit from TTS. Misread numbers, names, or medication instructions can cause real harm. Disclosure and consent matter whenever a voice resembles a real person.
How it works (plain)
A typical neural pipeline:
- Text frontend: normalize numbers/dates; sometimes convert graphemes → phonemes.
- Acoustic model: predict speech features (e.g., Mel spectrogram) or discrete audio tokens from text.
- Vocoder / waveform model: turn features/tokens into a playable waveform.
- Optional controls: speaking rate, pitch, speaker ID embedding, language ID (multi-speaker / multi-language recipes).
ESPnet’s TTS recipe documents stages from data prep through training, decoding, and evaluation—showing TTS is an engineering pipeline, not a single magic call.
CS224S homework explicitly includes synthesizing audio alongside transcripts—synthesis is a first-class spoken-language skill.
Everyday example
E-book read-aloud helps a tired eyes day. A bank that “verifies you by voice alone” is a different story: synthetic speech can fool casual listeners—use out-of-band verification for money and identity.
Try it
Generate TTS for a short paragraph you wrote (any commercial or open tool). Rate: (1) clarity, (2) name pronunciation, (3) number reading. Note one failure. Do not attempt to imitate a real person’s voice without consent.
Myths
- ⚠️ Myth: Natural voice proves a human is speaking.
- ✓ Reality: Synthetic speech can be highly realistic.
- ⚠️ Myth: TTS errors are only funny.
- ✓ Reality: Number and medication misreads can be harmful.
- ⚠️ Myth: Multi-speaker training equals ethical voice products.
- ✓ Reality: Training data consent and user disclosure are separate requirements.
Sources
- ESPnet TTS recipe (tts1): https://espnet.github.io/espnet/recipe/tts1.html ↗
- ESPnet TTS realtime demo notebook index: https://espnet.github.io/espnet/notebook/ ↗
- Stanford CS224S: https://web.stanford.edu/class/cs224s/index.html ↗
- Jurafsky & Martin TTS chapter (SLP3): https://web.stanford.edu/~jurafsky/slp3/ ↗
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
