COURSE 08L0100% FREE
Verified 2026-08-10

Multimodality

**Multimodal** systems handle more than one kind of signal—text plus images, audio, video, or other sensors.

What it is

Multimodal systems handle more than one kind of signal—text plus images, audio, video, or other sensors.

Why it matters

Real life is not only text. Accessibility tools, document apps that read screenshots, and creative suites increasingly combine modes—and inherit each mode’s failure modes (Course 29 deepfake literacy).

How it works (plain)

Often: encode image/audio into vectors the language model can attend to, then generate text (or the reverse for image generation). Training aligns those spaces with paired data.

Everyday example

“Explain this chart screenshot” or “describe this photo for alt text.”

Try it

Give an image-capable assistant a chart and ask for the trend—then verify the numbers yourself.

Myths

⚠️ Myth: If it can see, it can measure perfectly.
✓ Reality: OCR and chart-reading still err; verify critical figures.
⚠️ Myth: Multimodal means safe from deepfakes.
✓ Reality: Generative multimodal tools also create synthetic media.

Sources