AudioPod AI
  • Pricing

Stem Splitter

Getting a clean separation

The recording you feed in decides most of the result. What bleeds and why, why an MP3 of an MP3 is the worst case, and when a six-part split is worse than a four.

Lesson 3 · core · 7 min read

Open Stem Splitter

It is tempting to treat separation as an undo button for mixing. It is not. A finished song is one waveform, and the parts were added together in a way that cannot be reversed — what a separator does is estimate what each part probably sounded like, from what it can still hear of it. Everything in this lesson follows from that one fact.

The source decides the result

The largest single factor in how a split turns out is not a setting. It is the file you uploaded. Lossy compression works by discarding detail your ears are unlikely to miss — and a good deal of that discarded detail is exactly what the separation was going to use to tell two sources apart.

Write this
Not this
Where the file came from
The original file you bought, ripped, or exported yourself
Ripped from a video, or downloaded from a re-upload of a re-upload
Encoding history
Encoded once, from the master, or never encoded at all
Encoded, decoded and re-encoded several times over
The recording itself
A studio mix with the instruments spread across the stereo field
A live take with crowd noise, or a mono transfer of an old record
The master
Dynamic, with space between the loud parts and the quiet ones
Crushed as loud as it will go, everything squashed together

Where the file came from

Write this:
The original file you bought, ripped, or exported yourself
Not this:
Ripped from a video, or downloaded from a re-upload of a re-upload

Encoding history

Write this:
Encoded once, from the master, or never encoded at all
Not this:
Encoded, decoded and re-encoded several times over

The recording itself

Write this:
A studio mix with the instruments spread across the stereo field
Not this:
A live take with crowd noise, or a mono transfer of an old record

The master

Write this:
Dynamic, with space between the loud parts and the quiet ones
Not this:
Crushed as loud as it will go, everything squashed together

Careful

An MP3 of an MP3 is the worst realistic case. Each encode throws away a different slice of the top end, and the second pass throws away detail that was already damaged. You cannot get it back, and no mode or preset in this tool will make up for it — so if there is any chance of finding a cleaner copy, that is a better use of five minutes than re-running the split.

What bleeds, and why

Bleed — hearing a bit of one part inside another — is not a bug that will be fixed in a later version. It is what happens where two sources genuinely occupy the same frequencies at the same moment, and the separation has to attribute that shared energy to one of them.

Where you hear it
Why it happens
Cymbals in the vocal
Both are broad, bright and noisy rather than pitched. Sibilance and a crash share the same region, and there is often no clean line between them.
Kick drum in the bass
They live in the same low octaves and usually play on the same beats. A bass note under a kick hit is one blob of low-frequency energy.
Reverb and room in the wrong part
Reverb was printed into the mix as part of the sound. The tail of a vocal reverb is not clearly a vocal, so it tends to land in whatever part catches the leftovers.
Doubled or harmonised vocals
Backing lines that move with the lead share its timbre and timing. A separator splitting lead from backing is drawing a line the recording does not clearly contain.
Distorted guitar smeared everywhere
Distortion spreads a single instrument across the whole spectrum, so it overlaps almost everything else at once.

The pairs that bleed most often, and what they have in common.

Note

Almost every split has one weak file, and it is usually the one that catches whatever is left over. That is not the split failing — it is the leftovers going somewhere. Judge the job on the part you actually needed.

When six parts is worse than four

More parts sounds strictly better and is not, for a reason worth stating plainly: a part that is not in the recording does not come back empty. It comes back as a file with something else in it.

Run a Detailed split on a guitar-band song with no piano and you get a piano file. It will contain something — sustained notes from a synth pad, the harmonic body of a guitar chord, whatever most resembled a piano. Meanwhile those sounds have been taken OUT of the part they belonged in, so you have not just gained a useless file, you have degraded a useful one.

  • Ask what is actually playing before choosing a depth. Band for most songs; Detailed when there really is a distinct piano and guitar to find.
  • Karaoke is the cleanest split available, because two-way separation has the least to get wrong. If all you need is “with vocal” and “without”, take it.
  • The same trap applies to naming instruments: picking a saxophone from the catalogue when the song has no saxophone gives you a saxophone file, not an empty one.
  • If a part bled, go shallower rather than deeper. Bleed means you asked for a distinction the recording does not support, and more parts asks for more of them.

Tip

When you are unsure, run the shallower split first. It costs exactly the same as the deep one, because you are billed for the length of the song rather than the number of parts — so the only thing you lose by starting simple is a minute, and what you gain is a clean baseline to judge the deeper attempt against.

What a good result sounds like

Set the expectation correctly and you will stop re-running jobs that already worked. A good stem is usable, not pristine. It will be quieter and thinner than the same part sounded inside the mix, because the loudness of a master comes partly from everything sitting together. It will have some faint bleed at the edges. It may sound slightly watery under headphones on the quietest passages.

None of that is a failed split. What a failed split sounds like is different and unmistakable: the part you asked for is largely absent, or another instrument is clearly present at full volume. That is the result worth re-running — usually with a better source file, or a shallower split.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonSplit, Isolate, or ExcludeNext lessonWhat to do with stems

All Stem Splitter lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent