Video understanding overview
Models that interpret video: action recognition, tracking, temporal event detection, video captioning—not only single frames.
What it is
Models that interpret video: action recognition, tracking, temporal event detection, video captioning—not only single frames.
Why it matters
Security, sports, robotics, and accessibility tools use video understanding—with privacy and bias stakes.
How it works (plain)
Sample frames/clips → encode space+time → classify or caption. Compute costs rise vs images. Privacy: faces and locations need policy.
Try it
List one helpful and one harmful deployment of “always-on cameras + AI.”
Myths
- ⚠️ Myth: More cameras always mean more safety.
- ✓ Reality: Without governance, they mean more surveillance risk.
Sources
- Course 12 vision overview; Course 27 embodied; Course 20 society
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
