COURSE 15L1100% FREE
Verified 2026-08-14

Masked LMs and BERT

**BERT** (Bidirectional Encoder Representations from Transformers) pretrains a Transformer **encoder** on unlabeled text with **masked language modeling** and next-sentence-style objectives, then fine-tunes with a small task head for cla...

What it is

BERT (Bidirectional Encoder Representations from Transformers) pretrains a Transformer encoder on unlabeled text with masked language modeling and next-sentence-style objectives, then fine-tunes with a small task head for classification, QA, and more. Devlin et al. (NAACL 2019) showed one architecture + fine-tuning could hit strong results across many NLP tasks without heavy task-specific redesigns.

<!-- IMAGE: sentence with [MASK] tokens → bidirectional Transformer → predictions -->

HIGH PRIORITYDIAGRAM
◷ IN PRODUCTION

Visual Spec & Architecture Diagram

BERT MLM diagram: input sentence with [MASK] token; Transformer encoder; output distribution over vocab for the mask position showing top-3 fake candidates. Second small panel: NSP or sentence-pair optional. Label 'bidirectional context'.

Educational Focus: Masked LM intuition is the chapter's keystone visual.

Why it matters

BERT popularized “pretrain once, fine-tune many times” for understanding tasks (not chat decoding). It is still the conceptual parent of many encoders used for classification, NER, and rerankers—even in an LLM-first world. SLP3’s masked LM chapter and CS224N’s Transformer units teach the same idea.

How it works (plain)

  1. Hide some tokens in a sentence ([MASK]).
  2. Train the model to predict them using left and right context.
  3. After pretraining, attach a simple head and fine-tune on labeled data.
  4. Use the encoder outputs for classification ([CLS]), spans (QA/NER), or sentence pairs.

Everyday example

Fine-tune an encoder to flag “billing” vs “technical” tickets. You often need far less labeled data than training from scratch—and more control than free-form chat.

Try it

Read the BERT abstract. Write: (1) what “bidirectional” buys you vs left-to-right GPT-style LMs, (2) one task you’d fine-tune vs one you’d leave to a generative LLM.

Myths

⚠️ Myth: LLMs made BERT irrelevant.
✓ Reality: Encoders still power classifiers, retrievers, and NER at lower cost.
⚠️ Myth: Masking is just “fill in the blank for fun.”
✓ Reality: It is a pretraining objective that builds transferable representations.
⚠️ Myth: Fine-tuning never needs good labels.
✓ Reality: Garbage labels still produce garbage heads.

Sources