COURSE 13L1100% FREE
Verified 2026-08-14

Speech translation: cascades to end-to-end

**Speech translation (ST)** converts speech in one language into text (or speech) in another. For three decades the field moved from **cascades** (ASR then MT) toward **tighter coupling** and **end-to-end** models that map speech → targe...

What it is

Speech translation (ST) converts speech in one language into text (or speech) in another. For three decades the field moved from cascades (ASR then MT) toward tighter coupling and end-to-end models that map speech → target text directly. Sperber & Paulik (ACL 2020) survey that history and warn that many “end-to-end” systems still make compromises because labeled ST data is scarce.

<!-- IMAGE: cascade mic→ASR text→MT vs end-to-end mic→target text -->

HIGH PRIORITYDIAGRAM / COMPARISON
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

Two parallel architectures: CASCADE = Speech → ASR → text → MT → text (with error snowball icons between stages); E2E = Speech → Speech Translation model → target text (single model box). Bottom: pros/cons table as icons (latency, error compounding, data needs, interpretability).

Educational Focus: Cascade vs E2E is the chapter thesis—diagram locks the contrast.

Why it matters

Meetings, travel, accessibility, and multilingual support products depend on ST. Cascade errors compound (ASR mistake becomes MT mistake). End-to-end promises joint training—but data reality often reintroduces intermediate steps or multi-task tricks. Whisper-style systems also train X→en translation as a multitask goal alongside transcription.

How it works (plain)

Cascade: speech → ASR transcript → machine translation → (optional TTS). Clear modules; error propagation; needs good ASR *and* MT.

End-to-end: speech encoder → decoder in the target language. Joint optimization; needs parallel speech–translation data (rarer than ASR transcripts).

Hybrid / multitask reality: pretrain pieces, multi-task with ASR, use synthetic data, or weakly supervised translation hours (Whisper reports large X→en translation subsets).

ESPnet notebooks include speech translation demos and CMU assignments on SOTA ST models—good labs for seeing cascades vs E2E in practice.

Everyday example

Watching a talk in Spanish with English captions. A cascade might garble a name in ASR, then “helpfully” translate the wrong name. An E2E model might skip the explicit transcript—but can still hallucinate. Always verify critical entities.

Try it

Take one short utterance. Compare: (1) ASR then MT, vs (2) a direct ST mode if available. List errors that appeared only in the cascade vs only in direct ST.

Myths

⚠️ Myth: “End-to-end” on a slide means intermediate representations are gone.
✓ Reality: Sperber & Paulik argue many E2E systems still compromise under data scarcity—read claims carefully.
⚠️ Myth: Cascade is obsolete.
✓ Reality: Cascades remain strong, debuggable baselines—especially when you already have excellent ASR and MT.
⚠️ Myth: Fluent target text means correct meaning.
✓ Reality: Fluent mistranslation is worse than awkward honesty.

Sources