Vision-language models overview
**Vision-language models (VLMs)** connect images (and sometimes video) with text—captioning, visual Q&A, document screenshots, UI understanding.
What it is
Vision-language models (VLMs) connect images (and sometimes video) with text—captioning, visual Q&A, document screenshots, UI understanding.
Why it matters
Multimodal assistants are mainstream. They inherit vision errors *and* language hallucinations—verify critical reads.
How it works (plain)
Encode image → fuse with language model → generate or classify. Training uses paired image-text data. OCR and chart reading remain brittle.
Everyday example
“What does this screenshot’s error message say?” then confirm by eye before changing production.
Try it
Ask a VLM to read a chart; check every number yourself.
Myths
- ⚠️ Myth: If it can see, it can audit.
- ✓ Reality: Object hallucination and OCR errors persist.
Sources
- Course 08 multimodality; Course 12 vision; Course 07 LMs
- Model cards (cite specifically)
