🎧 Listen to this article
On This Page
0%- What is sonic branding, and does a small channel actually need one?
- The do / don't table
- Rule by rule, with the reasoning
- 1. Five pieces, not fifty
- 2. Three seconds, hard limit
- 3. Ducking beats "just turn it down"
- 4. One voice, held
- 5. Master once, export per platform
- How do I build a bed from music I already like?
- Where the kit gets applied
- FAQ
- A 90-minute build plan
Most channels get a logo before they get a sound. Then every episode opens with whichever royalty-free track was at the top of the search results that week, and six months of output has no audible through-line. Sonic branding is the cheap correction: a small, fixed set of audio assets you reuse until you are personally sick of them — which is roughly the point at which the audience starts recognizing them.
TL;DR — five rules for a channel that sounds like itself:
- Build one audio kit — intro sting, outro, two music beds, one ad bed — and reuse it for a year.
- Keep the intro sting under three seconds.
- Duck music under speech instead of globally lowering it.
- Lock one narrator voice per show and never swap mid-season.
- Master once, then export per platform.
What is sonic branding, and does a small channel actually need one?
Sonic branding is the audio equivalent of a color palette: a sting, a bed, a voice, and a fixed set of levels you apply the same way every time. It is not a jingle contest. Nobody needs a composer on retainer.
The threshold for bothering isn't audience size — it's episode count. Recognition is a repetition effect, so a kit only earns its keep if the same three seconds play across dozens of uploads. If you plan to publish more than a dozen times on any channel, build the kit before episode five, when re-cutting your back catalogue is still a two-hour job rather than a two-week one.
The do / don't table
| Do | Don't | Why it matters |
|---|---|---|
| Build a fixed five-piece kit | Pick new music per episode | Repetition is the entire mechanism; novelty destroys it |
| Keep the sting ≤ 3 seconds | Open with a 15-second music bed | Platform drop-off is highest in the first 10 seconds |
| Duck the bed under dialogue | Set the bed low and leave it | Static beds either fight speech or vanish entirely |
| Use one narrator voice per show | Rotate voices between episodes | Voice is the strongest recognition cue you own |
| High-pass the bed around 200 Hz | Run the bed full-range | Clears room for vocal fundamentals without lowering the bed |
| Master once to a clean reference | Master separately per upload | Divergent masters make your catalogue sound inconsistent |
Version your kit (kit-v1, kit-v2) | Overwrite the old files | You will need to re-cut old episodes; you need the originals |
| Give ad reads their own bed | Reuse the editorial bed under sponsorships | Audiences deserve an audible tell; regulators increasingly agree |
| Leave 300 ms of silence after the sting | Hard-cut sting into first word | The gap reads as intentional; the collision reads as amateur |
| Cap the sting at −1 dBTP | Normalize the sting to 0 dBFS | Lossy encoders overshoot and clip on playback |
Rule by rule, with the reasoning
1. Five pieces, not fifty
The working kit is: an intro sting (2–3 s), an outro (5–10 s, ideally the same motif resolved), a neutral editorial bed, a second bed with slightly more energy for mid-roll or list segments, and one distinctly different ad bed. That's it. Generate them in a single session so they share key and tempo — AudioPod's music tools let you prompt variations from one seed idea, which is the fastest way to get five cues that sound related rather than five cues that sound unrelated.
One key and one tempo across the kit is the trick most creators skip. It means any two pieces can butt against each other in an edit without a jarring modulation, and it means you can loop a bed under a segment of any length.
2. Three seconds, hard limit
Every major short-form and podcast platform front-loads its retention curve. A long musical intro spends your most valuable seconds on something the returning listener already knows. Three seconds is enough for a motif; anything longer is self-indulgence you pay for in drop-off. If you want a longer musical moment, put it at the outro, where it plays over the credits and costs you nothing.
3. Ducking beats "just turn it down"
A bed at a fixed low level is wrong in both directions: too loud in the quiet passages, inaudible under an animated guest. Sidechain the bed to the dialogue track — roughly 200 ms attack, 400–600 ms release, 8–12 dB of reduction — so the music breathes around the words. In practice you want the bed sitting about 15–20 dB below dialogue during speech and rising to about 8–10 dB below in the gaps.
If you're editing in a browser, AudioPod's DAW handles the multitrack side; Descript and Adobe Podcast have their own ducking implementations, and Audacity can do it manually with envelope points if you don't mind the tedium.
One extra move: high-pass the bed at 200–300 Hz. You free up the range where male vocal fundamentals live without touching the bed's perceived volume, which means less ducking is needed in the first place.
4. One voice, held
If your show uses synthetic narration — for a faceless channel, a translated edition, or a segment host — pick one voice and stay with it. Voice is a stronger identity signal than any music cue, and swapping it reads to the listener as a change of show. AudioPod's text-to-speech supports custom voice models on the Creator tier and above, which is the durable answer: a model you control doesn't get deprecated out from under your back catalogue the way a shared stock voice can. Vendors like ElevenLabs, Murf and Play.ht offer their own cloning paths — the point is to commit to one and keep the reference audio archived.
5. Master once, export per platform
Produce one clean master, then render platform-specific exports from it rather than remastering each time. Published targets differ, and they move — treat this as a starting point and check the current vendor documentation before a launch:
| Destination | Commonly published target | Practical note |
|---|---|---|
| Podcast apps | around −16 LUFS (stereo) | The long-standing podcast convention |
| Music streaming | around −14 LUFS | Louder masters get turned down, not up |
| YouTube | around −14 LUFS | Same normalization logic as music services |
| Short-form video | around −14 LUFS | Assume phone speakers and heavy compression |
More detail on that lives in our loudness guide on the blog. The reason to master once is consistency: two masters made a month apart will diverge in tone even with identical settings, and listeners hear it as a production quality drop.
How do I build a bed from music I already like?
The honest answer has a legal half and a technical half.
The technical half is easy. AudioPod's stem splitter will pull an instrumental out of a finished track, and so will LALAL.AI or Moises. A vocals-removed instrumental is usually a better bed than the full mix, because the vocal range is exactly where your speech needs room.
The legal half is where creators get into trouble. Separating stems doesn't create a licence. If you don't own or haven't licensed the track, an instrumental version of it is still the track, and content-ID systems catch it. Use stem separation on music you own, music you've licensed for the use, or your own generated cues. For a channel kit, generated cues you own outright are simply the lower-risk path — that's the practical argument for tools like AudioPod's music generation, Suno, or Udio over lifting a bed from a commercial release.
Where the kit gets applied
Once the five pieces exist, application is mechanical:
- Podcast episodes — sting, 300 ms gap, cold open, bed under the intro read, ad bed at the mid-roll marker, outro. Podcast Studio handles the assembly.
- Short-form cuts — sting only if the clip runs over 30 seconds; below that it eats the hook.
- Video voiceover — bed only under the intro and outro, silent under the body. Music under continuous narration fatigues viewers faster than it engages them.
- Audiobooks — no bed at all inside chapters. Retailers reject beds under narration in most cases; keep the kit to the front and back matter only.
FAQ
How long should a sonic brand last before I refresh it? Longer than feels comfortable. A year is a reasonable minimum. If the kit is generating complaints rather than boredom, that's a signal; your own fatigue isn't.
Do I need different stings for each platform? No. One sting, exported at each platform's loudness target. Different stings defeat the purpose.
What file format should I archive the kit in? WAV at 48 kHz, 24-bit, plus the project file. Archive the un-mastered versions too — you will want to re-render at a different loudness target eventually.
Can I use the same bed under sponsored segments? You can, but don't. A distinct ad bed gives listeners an audible boundary between editorial and paid content, which is good practice and increasingly an expectation of disclosure rules in several markets.
How long will my generated cues stay available for re-download? On AudioPod, generated files are retained by tier — 1 year on the free Basic plan, 2 years on Creator, 3 on Pro, 5 on Studio — and opening or downloading a file in the last 30 days extends its window. Job history stays in the dashboard regardless, so any project can be re-rendered in a click. Details are on /pricing.
What's the minimum spend to get a usable kit? Zero, to try it. The free tier includes every tool with 1,000 monthly credits, which is enough to generate and audition a handful of cues; a full kit with iterations is comfortably inside the Creator tier at $20/month. See /pricing for the full ladder.
A 90-minute build plan
- 0–20 min — Write down three adjectives for the show. Generate 8–10 short cues against them in one session, same key and tempo.
- 20–40 min — Shortlist three. Cut a 3-second sting and a 10-second outro from the strongest one.
- 40–60 min — Build the two editorial beds and the ad bed from the remaining two cues. High-pass each at 200 Hz.
- 60–75 min — Record or generate the standard intro read with your locked narrator voice. Set the ducking curve once and save it as a preset.
- 75–90 min — Master to your reference target, export per platform, and archive everything as
kit-v1.
The kit is not the creative work. It's the frame around the creative work, and its whole value is that you stop thinking about it after the first afternoon.

