Context windows and memory
A model’s **context window** is how much tokenized text it can consider at once—the working desk surface. It is **not** the same as human long-term memory.
What it is
A model’s context window is how much tokenized text it can consider at once—the working desk surface. It is not the same as human long-term memory.
Why it matters
Long chats get truncated. Huge PDFs may not fit. Products add retrieval or summaries (Course 09) when the desk is too small.
How it works (plain)
Everything in the window influences next-token prediction. Fall off the window and earlier details vanish unless the app stored them elsewhere (agent memory stores—Course 10).
Everyday example
You paste a 200-page contract; the tool silently only “sees” the beginning/end depending on design. Always check product limits.
Try it
Ask a chatbot what its stated context limit is (approximate). Then ask it to summarize a long pasted text and spot what it missed.
Myths
- ⚠️ Myth: Bigger windows mean the model remembers your life forever.
- ✓ Reality: Windows are per request/session machinery; long-term memory needs product features.
- ⚠️ Myth: If it fit once, it will use every token wisely.
- ✓ Reality: Attention can dilute; middle-of-context failures are a known practical issue in some settings.
Sources
- Hugging Face LLM course: https://huggingface.co/learn/llm-course/ ↗
- ANN live: https://www.ainerdnetwork.com/learn/context-windows-and-memory ↗
- Provider docs for context limits (cite the specific model card/docs you use)
