Evaluation of generative media
How to judge image/audio/video generators: human preference, task fitness, safety filters, and automated proxies with known limits.
What it is
How to judge image/audio/video generators: human preference, task fitness, safety filters, and automated proxies with known limits.
Why it matters
Pretty samples ≠ product quality. Eval prevents shipping brittle styles or unsafe failure modes.
How it works (plain)
Define the job (product shot? concept art?) → golden prompts → side-by-side prefs → check text fidelity / artifacts → safety cases → cost/latency.
Try it
Build a 10-prompt golden set for one use case; score two tools blind.
Myths
- ⚠️ Myth: FID alone picks the best model for your brand.
- ✓ Reality: Distribution metrics ≠ brand fit.
Sources
- Course 18 evaluation; Course 11 generative chapters
- Model card metrics sections (cite specifically)
