Machine translation basics
**Machine translation (MT)** maps text from one language to another. Modern MT is largely neural encoder–decoder / Transformer seq2seq. SLP3 includes a dedicated MT chapter; BART-style denoising pretraining also transfers to translation ...
What it is
Machine translation (MT) maps text from one language to another. Modern MT is largely neural encoder–decoder / Transformer seq2seq. SLP3 includes a dedicated MT chapter; BART-style denoising pretraining also transfers to translation settings (Lewis et al., ACL 2020).
<!-- IMAGE: source sentence → encoder-decoder → target sentence -->
Visual Spec & Architecture Diagram
Classic encoder–decoder MT: source sentence tokens into Encoder stack; context vectors; Decoder generating target tokens one-by-one with attention arrows from decoder step to source positions. Labels: 'source language', 'target language', 'attention'. Toy pair EN→ES 'the cat → el gato'.
Why it matters
Global business, support docs, and accessibility depend on MT. Errors in legal/medical domains can be costly. Fluent output can still be wrong—fluency ≠ fidelity.
How it works (plain)
- Encode source tokens into hidden states.
- Decode target tokens (autoregressive).
- Optionally constrain terminology with glossaries.
- Evaluate with automatic metrics and bilingual human review for critical content.
- For speech inputs, see Course 13 speech translation (cascades vs end-to-end).
CS224N treats Transformers and LLMs as core skills—the same stack powers modern MT systems students fine-tune or evaluate.
Everyday example
Translating a help center article, then having a bilingual human spot-check refund and safety pages.
Try it
Translate a paragraph round-trip (A→B→A). Note what drifted: names, negation, units, formality.
Myths
- ⚠️ Myth: Perfect fluency means perfect fidelity.
- ✓ Reality: Fluent mistranslation is dangerous.
- ⚠️ Myth: One BLEU point always equals better UX.
- ✓ Reality: Metrics are proxies—humans decide for high stakes.
- ⚠️ Myth: Multilingual LLMs remove the need for MT evaluation.
- ✓ Reality: You still need test sets per language pair and domain.
Sources
- SLP3 MT chapter: https://web.stanford.edu/~jurafsky/slp3/ ↗
- BART paper (seq2seq pretraining, MT gains reported): https://aclanthology.org/2020.acl-main.703/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
- Speech translation survey (adjacent): https://aclanthology.org/2020.acl-main.661/ ↗
