COURSE 13L1100% FREE
Verified 2026-08-14

Text-to-speech basics

**Text-to-speech (TTS)** turns written text into spoken audio. Modern neural TTS can sound natural. The same progress raises impersonation risk—this chapter teaches **how TTS systems are built and evaluated**, not how to clone a person w...

What it is

Text-to-speech (TTS) turns written text into spoken audio. Modern neural TTS can sound natural. The same progress raises impersonation risk—this chapter teaches how TTS systems are built and evaluated, not how to clone a person without consent. For scam literacy, see Course 13 unit-05 and Course 29.

<!-- IMAGE: text → linguistic features → spectrogram → waveform -->

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

TTS pipeline: Text → Text normalization → Linguistic features / phonemes → Acoustic model (mel spectrogram) → Vocoder (waveform) → Speaker. Show mel heatmap between acoustic model and vocoder. Labels: 'duration', 'pitch/prosody', 'vocoder reconstructs audio'. Optional dual path: concatenative vs neural TTS as faded alternatives.

Educational Focus: Mirrors ASR in reverse; mel→vocoder is the key intuition for modern TTS.
MEDIUM PRIORITYUI WIREFRAME
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

UI wireframe of a TTS playground: text box, sliders for speaking rate / pitch / emotion (neutral, happy, calm), speaker dropdown with fake names 'Speaker A/B', waveform preview. Labels call out 'prosody ≠ content'.

Educational Focus: Shows product surface of TTS controls without implying identity theft.

Why it matters

Accessibility (screen readers), IVR phone trees, education, and creative tools benefit from TTS. Misread numbers, names, or medication instructions can cause real harm. Disclosure and consent matter whenever a voice resembles a real person.

How it works (plain)

A typical neural pipeline:

  1. Text frontend: normalize numbers/dates; sometimes convert graphemes → phonemes.
  2. Acoustic model: predict speech features (e.g., Mel spectrogram) or discrete audio tokens from text.
  3. Vocoder / waveform model: turn features/tokens into a playable waveform.
  4. Optional controls: speaking rate, pitch, speaker ID embedding, language ID (multi-speaker / multi-language recipes).

ESPnet’s TTS recipe documents stages from data prep through training, decoding, and evaluation—showing TTS is an engineering pipeline, not a single magic call.

CS224S homework explicitly includes synthesizing audio alongside transcripts—synthesis is a first-class spoken-language skill.

Everyday example

E-book read-aloud helps a tired eyes day. A bank that “verifies you by voice alone” is a different story: synthetic speech can fool casual listeners—use out-of-band verification for money and identity.

Try it

Generate TTS for a short paragraph you wrote (any commercial or open tool). Rate: (1) clarity, (2) name pronunciation, (3) number reading. Note one failure. Do not attempt to imitate a real person’s voice without consent.

Myths

⚠️ Myth: Natural voice proves a human is speaking.
✓ Reality: Synthetic speech can be highly realistic.
⚠️ Myth: TTS errors are only funny.
✓ Reality: Number and medication misreads can be harmful.
⚠️ Myth: Multi-speaker training equals ethical voice products.
✓ Reality: Training data consent and user disclosure are separate requirements.

Sources