COURSE 03L1100% FREE
Verified 2026-08-10

Unsupervised and self-supervised learning

**Unsupervised learning** finds structure **without** human answer keys: clusters, compression, anomalies. **Self-supervised learning** creates its own practice problems from raw data—like hiding a word and predicting it. Modern language...

What it is

Unsupervised learning finds structure without human answer keys: clusters, compression, anomalies.

Self-supervised learning creates its own practice problems from raw data—like hiding a word and predicting it. Modern language models lean hard on this idea.

Why it matters

Labels are expensive. The internet (and companies’ logs) contain oceans of unlabeled text, images, and audio. Self-supervision is how foundation models soak that up.

How it works (plain)

  • Unsupervised: “group similar customers” or “compress this table”
  • Self-supervised: “mask part of the input; predict the missing piece”

After self-supervised pretraining, teams often add a smaller supervised step for a specific product task.

Everyday example

Topic clusters in support tickets (unsupervised). A model that learned English by predicting hidden words, then fine-tuned to classify refund requests (self-supervised → supervised).

Try it

Take 12 news headlines. Without software, pile them into 3 piles by theme. That is clustering by hand.

Myths

⚠️ Myth: Unsupervised means no human influence.
✓ Reality: Feature choices, similarity metrics, and data collection still encode decisions.
⚠️ Myth: Self-supervised equals unsupervised forever.
✓ Reality: Products usually add supervised or preference stages later.

Sources