Audio event detection
**Audio event detection (AED)** / **sound event detection (SED)** finds non-speech (or mixed) sounds in audio: alarms, glass break, dog bark, machine faults, sirens—not “what words were spoken,” but “what happened.” <!-- IMAGE: waveform ...
What it is
Audio event detection (AED) / sound event detection (SED) finds non-speech (or mixed) sounds in audio: alarms, glass break, dog bark, machine faults, sirens—not “what words were spoken,” but “what happened.”
<!-- IMAGE: waveform with event tags glass_break / alarm / speech -->
Visual Spec & Architecture Diagram
Sound Event Detection strip: waveform + parallel label tracks for 'dog bark', 'siren', 'speech', 'glass break' as colored bars with onset/offset. Side: 'frame-level classifier → postprocess'. Fake urban scene labels only.
Why it matters
Accessibility (environmental awareness), industrial monitoring, and safety alerting benefit. Always-on microphones raise privacy stakes: edge inference helps, but logs and uploads still need policy. This is a sibling of ASR in the audio family—not a replacement.
How it works (plain)
- Cut audio into short windows (or use a streaming buffer).
- Compute features or embeddings (filterbanks, pretrained audio encoders).
- Classify each window or detect event boundaries (onset/offset).
- Apply a threshold chosen for an acceptable false-alarm rate.
- Escalate high-impact alerts to a human.
Speech toolkits focus more on ASR/TTS/SE, but the evaluation mindset transfers: pick operating points, report precision/recall, and test in the real acoustic environment—not only quiet labs. CMU speech courses stress experimental practice alongside algorithms; the same discipline applies to event detectors.
Everyday example
A factory mic that should catch a bearing squeal. Too-low threshold → alarm fatigue. Too-high → missed failures. The threshold is a business decision, not a default.
Try it
List three events worth detecting at home or work, and three that would feel invasive if continuously monitored. For each “worth it” event, write the worst false alarm and the worst miss.
Myths
- ⚠️ Myth: Edge devices make privacy automatic.
- ✓ Reality: Local inference helps; cloud backups, analytics, and debug uploads still need rules.
- ⚠️ Myth: High accuracy on a public dataset guarantees field performance.
- ✓ Reality: Mic placement, background noise, and event rarity dominate.
- ⚠️ Myth: If ASR is solved, events are easy.
- ✓ Reality: Different labels, imbalance, and temporal structure.
Sources
- Course 13
audio-and-speech-modelsfamily map - ESPnet SE / related audio notebooks (enhancement adjacent): https://espnet.github.io/espnet/notebook/ ↗
- CMU 18-781 (experimental practice culture): https://courses.ece.cmu.edu/18781 ↗
- Course 19 privacy; Course 26 HAI
