Embeddings
An **embedding** turns a token, sentence, or document into a list of numbers (a vector) so a computer can compare “meaning-ish” similarity with math.
What it is
An embedding turns a token, sentence, or document into a list of numbers (a vector) so a computer can compare “meaning-ish” similarity with math.
Why it matters
Search, recommendations, and RAG (Course 09) often start by embedding text and finding nearest neighbors.
How it works (plain)
Items that appear in similar contexts get vectors that sit nearer in space. “King” and “queen” may end up closer than “king” and “broccoli”—depending on the model and data.
Everyday example
“Find docs like this ticket” in a help desk search box is often embedding similarity under the hood.
Try it
Write three sentences: two about dogs, one about finance. You would expect the dog sentences to embed closer together in a decent model.
Myths
- ⚠️ Myth: Embedding distance equals true meaning.
- ✓ Reality: It reflects training patterns; biases and quirks transfer.
- ⚠️ Myth: One embedding model fits all languages and domains.
- ✓ Reality: Domain mismatch hurts retrieval quality.
Sources
- Hugging Face embeddings / sentence transformers docs: https://huggingface.co/blog/getting-started-with-embeddings ↗
- ANN live: https://www.ainerdnetwork.com/learn/embeddings ↗
- D2L representation learning chapters: https://www.d2l.ai/ ↗
