Masked LMs and BERT
**BERT** (Bidirectional Encoder Representations from Transformers) pretrains a Transformer **encoder** on unlabeled text with **masked language modeling** and next-sentence-style objectives, then fine-tunes with a small task head for cla...
What it is
BERT (Bidirectional Encoder Representations from Transformers) pretrains a Transformer encoder on unlabeled text with masked language modeling and next-sentence-style objectives, then fine-tunes with a small task head for classification, QA, and more. Devlin et al. (NAACL 2019) showed one architecture + fine-tuning could hit strong results across many NLP tasks without heavy task-specific redesigns.
<!-- IMAGE: sentence with [MASK] tokens → bidirectional Transformer → predictions -->
Visual Spec & Architecture Diagram
BERT MLM diagram: input sentence with [MASK] token; Transformer encoder; output distribution over vocab for the mask position showing top-3 fake candidates. Second small panel: NSP or sentence-pair optional. Label 'bidirectional context'.
Why it matters
BERT popularized “pretrain once, fine-tune many times” for understanding tasks (not chat decoding). It is still the conceptual parent of many encoders used for classification, NER, and rerankers—even in an LLM-first world. SLP3’s masked LM chapter and CS224N’s Transformer units teach the same idea.
How it works (plain)
- Hide some tokens in a sentence (
[MASK]). - Train the model to predict them using left and right context.
- After pretraining, attach a simple head and fine-tune on labeled data.
- Use the encoder outputs for classification (
[CLS]), spans (QA/NER), or sentence pairs.
Everyday example
Fine-tune an encoder to flag “billing” vs “technical” tickets. You often need far less labeled data than training from scratch—and more control than free-form chat.
Try it
Read the BERT abstract. Write: (1) what “bidirectional” buys you vs left-to-right GPT-style LMs, (2) one task you’d fine-tune vs one you’d leave to a generative LLM.
Myths
- ⚠️ Myth: LLMs made BERT irrelevant.
- ✓ Reality: Encoders still power classifiers, retrievers, and NER at lower cost.
- ⚠️ Myth: Masking is just “fill in the blank for fun.”
- ✓ Reality: It is a pretraining objective that builds transferable representations.
- ⚠️ Myth: Fine-tuning never needs good labels.
- ✓ Reality: Garbage labels still produce garbage heads.
Sources
- BERT paper: https://aclanthology.org/N19-1423/ ↗
- SLP3 masked LM chapter: https://web.stanford.edu/~jurafsky/slp3/ ↗
- Stanford CS224N: https://web.stanford.edu/class/cs224n/ ↗
