COURSE 12L1100% FREE
Verified 2026-08-10

Vision-language models overview

**Vision-language models (VLMs)** connect images (and sometimes video) with text—captioning, visual Q&A, document screenshots, UI understanding.

What it is

Vision-language models (VLMs) connect images (and sometimes video) with text—captioning, visual Q&A, document screenshots, UI understanding.

Why it matters

Multimodal assistants are mainstream. They inherit vision errors *and* language hallucinations—verify critical reads.

How it works (plain)

Encode image → fuse with language model → generate or classify. Training uses paired image-text data. OCR and chart reading remain brittle.

Everyday example

“What does this screenshot’s error message say?” then confirm by eye before changing production.

Try it

Ask a VLM to read a chart; check every number yourself.

Myths

⚠️ Myth: If it can see, it can audit.
✓ Reality: Object hallucination and OCR errors persist.

Sources

  • Course 08 multimodality; Course 12 vision; Course 07 LMs
  • Model cards (cite specifically)