COURSE 05L1100% FREE
Verified 2026-08-10

CNNs and computer vision

**CNNs (convolutional neural networks)** are neural nets designed for grid-like data such as images. They slide small filters across the image to detect local patterns—edges, textures—then build toward objects.

What it is

CNNs (convolutional neural networks) are neural nets designed for grid-like data such as images. They slide small filters across the image to detect local patterns—edges, textures—then build toward objects.

Why it matters

Modern photo tagging, medical imaging prototypes, warehouse scanners, and many video tools grew from CNN breakthroughs (with transformers now also competing/combining in vision).

How it works (plain)

A filter looks at a small patch. Many filters look for different patterns. Pooling/striding shrinks maps. Deeper layers combine simple patterns into complex ones.

Everyday example

Unlocking a phone with a face camera (with liveness checks in serious products). Sorting recyclables on a belt with a camera.

Try it

Zoom into a photo until you only see edges. That is closer to what early CNN layers emphasize.

Myths

⚠️ Myth: CNNs “see” like humans.
✓ Reality: They respond to statistical patterns and can be fooled by adversarial noise or weird contexts.
⚠️ Myth: Vision is solved.
✓ Reality: Robustness, rare cases, and fairness issues remain (Course 19).

Sources