Language models
A **language model** learns to guess the next piece of text. Given “The sky is ___,” it assigns high probability to words like “blue”—not because it looked outside, but because that pattern appeared often in training data. Large language...
What it is
A language model learns to guess the next piece of text. Given “The sky is ___,” it assigns high probability to words like “blue”—not because it looked outside, but because that pattern appeared often in training data.
Large language models (LLMs) are the same idea at huge scale, usually built with transformers, then adapted so chat feels helpful.
Why it matters
Chat, coding assistants, summarizers, and many “AI agents” are wrapped around next-token prediction. If you understand that core loop, the rest of Course 07 clicks.
How it works (plain)
- Split text into tokens (Course chapter on tokenization).
- During training, predict each next token; raise the score of the true next token.
- At use time, take your prompt, predict a token, append it, repeat.
There is no separate “truth module.” Helpfulness and correctness are emergent—and imperfect.
Everyday example
Autocomplete on your phone is a tiny cousin. Chatbots are a giant cousin with more training and chat formatting.
Try it
Start a sentence and write two different plausible endings. A language model is doing a statistical version of that choice, millions of times per reply.
Myths
- ⚠️ Myth: The model looks up a database of facts for every answer.
- ✓ Reality: Base LMs predict tokens; tools/RAG can add lookup (Course 09).
- ⚠️ Myth: Fluent answers are true answers.
- ✓ Reality: Fluency and truth diverge (hallucinations—Course 08).
Sources
- Hugging Face LLM Course (Ch.1): https://huggingface.co/learn/llm-course/en/chapter1/1 ↗
- Jurafsky & Martin, *Speech and Language Processing* (LM chapters)—https://web.stanford.edu/~jurafsky/slp3/ ↗
- ANN live: https://www.ainerdnetwork.com/learn/language-models ↗
