Computer vision overview
**Computer vision** teaches machines to interpret images and video: what is there, where it is, and sometimes what will happen next.
What it is
Computer vision teaches machines to interpret images and video: what is there, where it is, and sometimes what will happen next.
Why it matters
Vision powers photo apps, industrial inspection, medical imaging assistance, robotics perception—and surveillance risks. Pair with Course 19/29 for misuse literacy.
How it works (plain)
Pipelines often: pixels → features (classical or learned) → task head (classify, detect boxes, segment masks, caption). Modern systems lean on CNNs and vision transformers.
Everyday example
Phone camera detecting a QR code or suggesting a crop—narrow vision tasks, not “understanding the whole world.”
Try it
Pick a photo and list: classify, detect, segment, caption. Which task matches your need?
Myths
- ⚠️ Myth: Vision models see like humans.
- ✓ Reality: They optimize training objectives; adversarial patches and dataset gaps reveal differences.
- ⚠️ Myth: Higher accuracy on a benchmark means safe deployment.
- ✓ Reality: Domain shift and demographic performance gaps still bite.
Sources
- Course 05 cnns-and-vision
- CS231n-style curricula (verify current offering)
- NIST AI RMF for risk framing: https://www.nist.gov/itl/ai-risk-management-framework ↗
