AudioPod AI
  • Pricing

Speaker Separation

How many speakers to declare

The one control that decides whether a separation is right. When automatic is enough, when to declare an exact number, and how to recognise the two ways it goes wrong without telling you.

Lesson 2 · core · 7 min read

Open Speaker Separation

The speaker count is the only real decision this tool asks you to make, and it is genuinely hard, because the two ways it goes wrong both look like success. The job turns green. The tracks appear. The credits are spent. You find out by listening, or you never find out.

Automatic, or a number

“Auto detect” asks the separation to work out how many distinct voices are in the recording. Declaring a number instead tells it, and it will produce that many. Neither is safer in general. The right one depends entirely on how certain you are.

Write this
Not this
You know the number for certain
Declare it, and get exactly that many tracks
Leave it automatic and hope
People join or leave partway through
Leave it on “Auto detect” — a fixed count is wrong for most of the recording
Declare the number who appear at some point
A recording you did not make
Run automatic first, then declare a count if the result is wrong
Guess a number from the first minute
Two people, one microphone, one room
Declare two, which forces the split it would not commit to
Leave it on “Auto detect” after it merged them once
A crowd, or heavy crosstalk
Accept that this is the case the tool handles worst, and edit the source instead
Declare a high count and expect clean tracks

You know the number for certain

Write this:
Declare it, and get exactly that many tracks
Not this:
Leave it automatic and hope

People join or leave partway through

Write this:
Leave it on “Auto detect” — a fixed count is wrong for most of the recording
Not this:
Declare the number who appear at some point

A recording you did not make

Write this:
Run automatic first, then declare a count if the result is wrong
Not this:
Guess a number from the first minute

Two people, one microphone, one room

Write this:
Declare two, which forces the split it would not commit to
Not this:
Leave it on “Auto detect” after it merged them once

A crowd, or heavy crosstalk

Write this:
Accept that this is the case the tool handles worst, and edit the source instead
Not this:
Declare a high count and expect clean tracks

The menu offers 2 through 10, so 10 people is the ceiling. That is not an arbitrary cap so much as an honest edge: a recording with more voices than that almost always has them overlapping, and overlapping speech is the thing separation is worst at.

The two failures nothing tells you about

Both of these finish as a successful job. There is no warning, no badge, and no refund. The only detector is your own ear, which is why the first lesson asks you to play thirty seconds of every track before doing anything else.

What you hear
What to do
The job finished and said “Only one speaker detected”. It heard a single voice, so there was nothing to pull apart. Usually the clip really is one person, sometimes two people were recorded so closely that they read as one.
Check the recording has two genuinely distinct voices in it. Nothing was separated, so there is nothing to salvage from this run.
The job finished and said “No separated tracks”. It could not find speech it was confident enough to split.
Try a clearer recording. Heavy background music, a very short clip, or speech buried under noise are the usual reasons.
Two people came back on the same track. Their voices were close enough — in pitch, in accent, or because they were on one microphone in one room — that they clustered together.
Re-run with the exact count declared. Telling it there are two forces the split it would not otherwise commit to.
One person came back as two speakers. Their voice changed enough partway through to look like a second person — moving away from the microphone, a phone line dropping quality, laughing, shouting.
Re-run with the exact count declared. If it survives that, merge the two in the transcript by reassigning the stray lines.

How a separation disappoints you, and what to do about it.

Careful

A re-run is a new job at the full rate. On a five-minute clip that is 1,650 credits and no reason to hesitate; on a two-hour panel it is 39,600, which is the argument for testing your settings on a short excerpt of a long recording before you commit the whole thing.

Tip

Cut a two-minute excerpt from the middle of a long recording and separate that first. 660 credits buys you the answer to “does the count I think is right actually work here”, and the middle is where crosstalk lives — the opening minute is the least representative part of any conversation.

The transcript switch, and why it is free

“Include Speaker-Labeled Transcription” sits under the count and starts off. It adds a timestamped transcript with a speaker label on every line, alongside the separated audio. It costs nothing extra. The credit estimate is the same either way.

That matters more here than it looks, because the transcript is the fastest way to audit a separation you are unsure about. Reading who was credited with which line takes a minute; listening to two full tracks takes as long as the recording. If the transcript shows one person's lines split across two labels, you have found a problem the audio would have taken an hour to reveal.

Careful

Turn it on BEFORE the run. The tool offers no way to add a transcript to a job that has already finished, so getting one afterwards means separating the whole recording again and paying for it again. Since it costs nothing extra, the only reason to leave it off is that you are certain you will never want it.

When only a few lines are wrong

Not every mistake is worth a re-run. If the tracks are broadly right and a handful of lines landed on the wrong person, fix it in the transcript instead. Move a segment to a different speaker, or to a new one you add. This is how you fix a line that landed on the wrong person. That corrects the record without spending anything.

Note

What a transcript edit does NOT do is move audio. Reassigning a line changes who the transcript says spoke it; the audio for that moment stays on the track the separation put it on. When the audio itself is wrong, a re-run with a declared count is the only fix.

Once the tracks are right, the question becomes what to do with them — which is the next lesson, and the reason most people came here in the first place.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonYour first separationNext lessonWorking with the separated tracks

All Speaker Separation lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent