🎧 Listen to this article
On This Page
0%- The do/don't table
- Why capture flat when the enhanced version sounds better in the room?
- Clean once, at the top of the chain
- Cut before you master, not after
- Why does each platform need its own loudness render?
- Keep music separable
- Transcribe once, derive everything
- Where the tools fit
- FAQ
- The one-sentence version
Most creators record once and then re-record for every platform they publish to. That's backwards. A single well-captured session already contains everything the podcast feed, the YouTube upload, the Shorts cut, the TikTok clip, and the newsletter audiogram need — what's missing is a processing order that doesn't destroy the material on the way to each destination.
TL;DR — the six-platform rule:
- Capture flat. No compression, no noise gate, no "enhance" at the source.
- Clean once, at the top. Denoise and de-reverb before any edit, never per-cut.
- Cut before you master. Clips inherit the master, not the other way round.
- Master per platform, not once for all six.
- Transcribe the master, then derive every caption from that one transcript.
- Keep the source project, not just the exports.
The do/don't table
| Step | Do | Don't | Why it matters |
|---|---|---|---|
| Capture | Record flat WAV, 48 kHz / 24-bit, peaks around -12 dBFS | Enable the interface's built-in compressor or "voice enhance" | Baked-in processing can't be undone; you inherit its artifacts in all six exports |
| Cleanup | Denoise and de-reverb once, on the full-length master | Run a denoiser separately on each clip | Per-clip cleanup drifts — different noise profiles across cuts make the same voice sound like three people |
| Editing | Cut from the cleaned master timeline | Cut first, clean the fragments after | Fragments give the denoiser no room to learn a noise floor between phrases |
| Music | Keep music on its own stem, ducked -12 to -18 dB under speech | Bounce a mixed bed you can't separate later | A Shorts cut usually wants louder music than the podcast does |
| Loudness | Master a separate render per platform target | Ship one -16 LUFS file everywhere | Every platform normalizes; a single target means five of six get re-processed |
| Captions | Transcribe the master once, slice the transcript per cut | Auto-caption each platform upload natively | Native auto-captions disagree with each other and with your show notes |
| Archive | Keep the project and the stems | Keep only the MP3s | Re-cutting a viral moment 6 months later needs the source, not the export |
Everything below is the reasoning behind a row.
Why capture flat when the enhanced version sounds better in the room?
Because "sounds better in the room" is a decision made once, at the worst possible time — before you know what you're cutting. Interface-level compression and enhancement are one-way doors. If a guest leans in and clips, a flat recording gives you headroom to pull it back; an already-compressed recording gives you a squashed transient with the noise floor pumped up around it.
The practical target: 48 kHz, 24-bit, peaks landing near -12 dBFS, with no processing in the chain. That leaves roughly 12 dB of headroom for the loud laugh you didn't plan for. 24-bit is what makes this safe — the noise floor sits far enough down that a quiet-but-clean recording can be brought up later without hiss coming with it. Recording hot to "use the full range" is a 16-bit habit that stopped mattering years ago.
If your room is bad, fix the room, not the signal chain. A moving blanket behind the mic does more than any real-time enhancement, and it doesn't cost you your options.
Clean once, at the top of the chain
This is the row creators break most often, and it's the one that produces the weirdest results. The instinct is reasonable: you cut a 45-second TikTok clip, notice a hum, and run cleanup on that clip. Then you cut a different moment two weeks later, run cleanup again, and the two clips sound like they were recorded in different rooms — because as far as the processing was concerned, they were.
Denoising works by estimating what's noise and what's signal. Give it 40 minutes of continuous material and the estimate is stable across the whole session. Give it 45 seconds that happen to start mid-sentence and it has almost nothing to learn from. Run noise reduction once on the full master, then never think about it again — every downstream cut inherits the same, consistent treatment.
The same argument applies to de-reverb and to any repair work. Do it while you still have the longest possible view of the material.
Cut before you master, not after
A master is a decision about a destination. A cut is a decision about content. Making the destination decision first means every cut carries the wrong one.
The order that works:
- Cleaned full-length master (one file, no loudness target yet)
- Editorial cuts — the full episode, the YouTube version, the three short-form moments
- Per-cut mix adjustments (music level, pacing)
- Per-platform loudness render, last
Step 4 is the only one that's platform-specific, and it's cheap because it's last. Reversing 3 and 4 means re-rendering a whole tree of files every time a platform changes its normalization behavior — which they do, quietly, without a changelog.
Why does each platform need its own loudness render?
Because every platform normalizes on ingest, and they don't agree on the target. If you ship one file everywhere, exactly one platform leaves it alone and the rest apply gain to it — sometimes turning it down, which is fine, sometimes turning it up and revealing the noise floor you thought you'd handled.
As of late 2026, the commonly-cited targets look roughly like this — check each platform's current spec before you commit, since these move:
| Destination | Typical integrated loudness target | True peak ceiling |
|---|---|---|
| Podcast feeds (Apple, Spotify) | around -16 LUFS mono / -16 to -14 LUFS stereo | -1 dBTP |
| YouTube long-form | around -14 LUFS | -1 dBTP |
| Short-form vertical (Shorts, TikTok, Reels) | around -14 LUFS, often played louder in-app | -1 dBTP |
| Audiobook retail (ACX-style) | -23 to -18 dB RMS, per the retailer's published spec | -3 dBTP |
| Newsletter audiogram | match the platform it embeds to | -1 dBTP |
The short-form row deserves a note. Vertical-video feeds are consumed at arm's length on a phone speaker, often in a noisy environment, and they're competing with the clip before yours. Speech that sits comfortably in a podcast mix reads as thin and distant there. The fix isn't more limiting — it's a slightly more aggressive compression ratio and a music bed pulled up a few dB relative to the podcast version. That's a mix decision, which is why the stems from step 3 matter.
Keep music separable
If your music arrives as a bounced bed mixed into the voice track, every downstream decision about music level is already made. That's a problem specifically because short-form and long-form want different answers.
Two ways to stay flexible. The clean way: keep music on its own track from the start, whether it's licensed or generated in Music Studio, and duck it under speech at render time. The recovery way, when you inherited a mixed file with no stems: run it through a stem splitter to pull the music and voice apart, then rebuild. Separation is good enough now that this is a real option rather than a last resort — but it's still recovery. Recording the stems separately costs nothing and always beats reconstructing them.
For speech-and-music ducking, -12 to -18 dB under the voice is the working range. Closer to -12 for short-form where energy matters; closer to -18 for interview podcasts where intelligibility does.
Transcribe once, derive everything
The biggest hidden cost in multi-platform publishing isn't audio — it's text. Show notes, YouTube description, chapter markers, TikTok captions, the blog post, the newsletter blurb. Creators generate each of these separately and end up with six slightly different versions of what was said.
Transcribe the cleaned master once with speech-to-text, and treat that transcript as the single source. Every caption file is a slice of it. Every chapter marker is a timestamp in it. The blog post is an edit of it. When a quote is wrong in one place, it's wrong in one place.
This also solves the caption-drift problem. Native auto-captioning on each platform produces a different transcript of the same words — different punctuation, different handling of names, different guesses at technical terms. Uploading your own caption file from the master transcript makes all six agree, and it's the version that gets indexed.
Where the tools fit
A lot of creators assemble this pipeline out of one subscription per step — a cleanup tool, a transcription tool, a separation tool, a mastering tool. That works, and it's how most people start. The friction shows up in the handoffs: re-uploading a 40-minute WAV four times, reconciling four sets of timestamps, discovering that the transcription tool's output format doesn't match what the caption tool wants.
The alternative is running the chain in one place. AudioPod covers cleanup, separation, transcription, voiceover, music, and multi-track assembly in the DAW, which means the master stays put and each step reads from it. Named competitors do individual steps well — Descript is strong on transcript-driven editing, LALAL.AI and Moises focus on separation, Adobe Podcast on speech cleanup — and if one of those is already load-bearing in your workflow, keep it. The point of the do/don't table isn't the tool, it's the order.
On cost: AudioPod is free to start (1,000 credits/month, every tool included), with paid plans from $20/mo on Creator. Full breakdown on /pricing.
FAQ
Should I record video and audio separately? Yes, if the audio matters. Camera audio is a reference track for syncing, not a deliverable. Record to a dedicated mic and interface, sync in post. The exception is a phone-only workflow where there's nothing to sync to — then get the phone close and treat the room acoustics as the whole game.
How do I handle a guest who recorded on a bad mic? Clean their track separately from yours — different source, different noise profile, so this is the one legitimate exception to "clean once." Match loudness between the two tracks after cleanup, not before. If the guest audio is unusable, a voice changer can't rescue a mangled source; re-record if the content justifies it.
What's the minimum viable version of this workflow? Capture flat → denoise the master → cut → render at -16 LUFS for podcast and -14 LUFS for everything video. That's four steps and it gets you 80% of the benefit. Add the transcript-as-source-of-truth step next; it saves the most time.
Do I need to re-render when a platform changes its loudness target? Only for that platform, and only if you kept the pre-master. This is the entire argument for the cut-before-master order. If you kept only the exports, you're re-processing an already-processed file, which compounds artifacts.
How long should I keep the source project? Longer than you think — the clip that goes viral is usually old. AudioPod retention runs 1 year on Basic, 2 years on Creator, 3 on Pro, 5 on Studio, and indefinite on Enterprise. Job history stays in the dashboard past the window, so a project can be re-rendered in one click, and anything you've downloaded in the last 30 days has its window extended automatically.
Can I turn a podcast episode into an audiobook chapter? If the content is structured for it, yes — but it's a rewrite, not a re-render. Spoken conversation and narrated prose have different rhythms. Audiobook Studio handles the narration side once you have the text; the editorial work of converting a transcript into prose is the actual job.
The one-sentence version
Process in the order that keeps your options open the longest: capture flat, clean once at the top, cut, and only then decide where it's going. Everything expensive in multi-platform publishing comes from making a destination decision before you had to. More workflow write-ups on the blog.

