Pretraining and fine-tuning
**Pretraining** teaches a model general patterns from huge corpora (often self-supervised next-token prediction). **Fine-tuning** continues training on narrower data so the model behaves better for a product: chat style, domain language,...
What it is
Pretraining teaches a model general patterns from huge corpora (often self-supervised next-token prediction).
Fine-tuning continues training on narrower data so the model behaves better for a product: chat style, domain language, tool formats, safer refusals.
Why it matters
You rarely train a giant LM from scratch. You adapt one. Understanding the stages explains why models know “a bit of everything,” then suddenly speak like a customer-support agent.
How it works (plain)
- Pretrain on broad text
- Optionally instruction-tune on examples of following directions
- Optionally preference-tune with human or AI feedback (Course 19)
- Ship with system prompts + tools + RAG
Everyday example
A general model fine-tuned on your company’s help articles and tone guide.
Try it
Write three instruction examples for a “polite librarian” assistant. That is the seed of an instruction set.
Myths
- ⚠️ Myth: Fine-tuning always beats good prompting + RAG.
- ✓ Reality: It depends on cost, update frequency, and risk; many teams mix approaches.
- ⚠️ Myth: Fine-tuning inserts a perfect knowledge database.
- ✓ Reality: It shifts behavior and weights; facts can still be wrong or stale.
Sources
- Hugging Face LLM course: https://huggingface.co/learn/llm-course/ ↗
- ANN live: https://www.ainerdnetwork.com/learn/pretraining-and-finetuning ↗
- Vaswani et al. 2017 (architecture foundation): https://arxiv.org/abs/1706.03762 ↗
