Language Models and Transformers
The modern engine of AI: transformers, multi-head self-attention, tokenization, embeddings, pretraining, fine-tuning, and Mixture of Experts (MoE).
Course Syllabus & Units
Lm Basics
3 lessonsEmbeddings
An **embedding** turns a token, sentence, or document into a list of numbers (a vector) so a computer can compare “meaning-ish” similarity with math.
Language models
A **language model** learns to guess the next piece of text. Given “The sky is ___,” it assigns high probability to words like “blue”—not because it looked outside, but because that pattern appeared often in training data. Large language...
Perplexity intuition
**Perplexity** summarizes how surprised a language model is by a text set—lower usually means better next-token prediction on that set.
Tokens
2 lessonsTokenization
Models do not read English letters the way you do. **Tokenization** splits text into **tokens**—chunks like whole words, word pieces, or punctuation—that the model scores.
Tokenizer multilingual issues
How tokenizers treat different languages unevenly—some languages need more tokens for the same meaning, raising cost and changing quality.
Transformers
4 lessonsPositional encodings
Transformers treat tokens as a set unless you tell them **order**. **Positional encodings** (or positional embeddings) inject “this is word 1, this is word 2…” so grammar and sequence make sense.
Query, key, value attention
Inside transformers, **attention** uses three roles for each token: a **query** (“what am I looking for?”), **keys** (“what do I contain?”), and **values** (“what content do I pass along?”).
Self-attention
**Self-attention** lets each token in a sequence build a weighted mix of information from other tokens in that same sequence. “Self” means the sequence attends to itself.
Transformers
A **transformer** is the neural network design behind most modern language models. Its key trick is **attention**: each token can look at other tokens and decide what matters for the next prediction—without reading the sequence only left...
Context
2 lessonsContext windows and memory
A model’s **context window** is how much tokenized text it can consider at once—the working desk surface. It is **not** the same as human long-term memory.
Long-context methods
Techniques that help models handle **longer inputs**—bigger context windows, smarter position schemes, retrieval instead of stuffing, and hierarchical summaries.
Training Stages
3 lessonsInstruction tuning and chat templates
**Instruction tuning** trains models to follow prompts/instructions. **Chat templates** format roles (system/user/assistant) the way a model expects.
LoRA and parameter-efficient finetuning
**Parameter-efficient finetuning (PEFT)** methods like **LoRA** adapt large models by training small adapter matrices instead of all weights.
Pretraining and fine-tuning
**Pretraining** teaches a model general patterns from huge corpora (often self-supervised next-token prediction). **Fine-tuning** continues training on narrower data so the model behaves better for a product: chat style, domain language,...
Decoding
2 lessonsSampling: temperature and top-p
After a model scores possible next tokens, **decoding** chooses among them. **Temperature** and **top-p** (nucleus sampling) are common knobs that trade boredom for creativity—and sometimes for chaos.
Speculative decoding overview
A serving trick: a **smaller draft model** proposes several tokens; the **large model** checks them in parallel and accepts a prefix—speeding inference without changing the big model’s distribution (when done correctly).
