Audio generation overview
Models that generate music, sound effects, or speech audio from prompts or references—sibling to TTS but broader.
What it is
Models that generate music, sound effects, or speech audio from prompts or references—sibling to TTS but broader.
Why it matters
Creative leverage is high; copyright, consent, and scam misuse (voice) are real. Pair with Course 13 and Course 29 scam literacy.
How it works (plain)
Text/audio conditioning → generative model (often diffusion/codec LMs) → waveform. Check licenses before commercial use.
Try it
Generate a short SFX bed for a personal project; document the tool’s license terms in one sentence.
Myths
- ⚠️ Myth: “AI music” is always copyright-free.
- ✓ Reality: Terms and training-data disputes vary—read the contract.
Sources
- Course 13 audio-and-speech-models; Course 11 generative overview
- Provider audio ToS (primary)
