AudioPod AI
  • Pricing

Speaker Separation

Your first separation

A recording of two people in, two clean tracks out. What Speaker Separation costs, what it hands back, and why the first run should leave every setting alone.

Lesson 1 · start · 4 min read

Open Speaker Separation

Speaker Separation takes one recording with several people in it and gives you back one audio track per person. That is the whole idea, and this lesson runs it once end to end on the easiest case there is: two people, talking in turn.

Note

There is no plan gate on this tool. Every account can run it, including a free one — the only limit is credits. That is worth saying plainly because most of the studio is not like that, and readers assume a lock that is not there.

Bring the audio in

Tab
What it takes
Upload
Drop in an audio or video file from your machine. The line above the dropzone shows the real size and length ceiling for your plan, read from the server rather than guessed.
URL
Paste a YouTube, Vimeo or direct audio or video link and it is fetched for you. For anything else, download the file first and upload it.

The two input tabs.

Up to 3 files at a time on the upload tab, each processed as its own job. How large and how long a single file may be depends on your plan, and the tool prints your real figures on the line above the dropzone. Do not go looking for those numbers in a guide: they are read from the server for your account, which is the only place they are ever correct.

Leave the settings alone

There are two controls, and on a first run you want neither of them. The speaker count starts on “Auto detect” and will work out how many people it can hear. The transcript switch starts off. Run it as it is, hear what comes back, and change one thing at a time after that.

Tip

The credit estimate appears beside the input tabs once the length is known. 330 credits per minute of audio you put in — 3,300 credits for ten minutes, 19,800 for an hour, however many people are in it.

What comes back

What
Detail
One track per speaker
Numbered Speaker 1, Speaker 2 and so on, each with its own waveform player. One track carries one person's voice and nobody else's.
Played and downloaded one at a time
Each track has its own play and download control. There is no zip and no download-all — take the speakers you need individually.
The optional transcript
Only when you asked for it before the run. It reads in the page, downloads as TXT or JSON, and can be corrected by hand.

The output of a finished job.

Careful

Each track has its own play and download control. There is no zip and no download-all — take the speakers you need individually. If you need all of them, that is several clicks rather than one, and it is worth knowing before you plan a workflow around it.

Check that it worked

  1. 01

    Count the tracks

    Two people should give you two. A third track on a two-person recording means one person was heard as two, which the next lesson covers.

  2. 02

    Play the start of each one

    You are listening for one voice and one voice only. A second voice under the first is the failure that matters most, and it is the one nothing on screen will tell you about.

  3. 03

    Skip to the middle and play again

    A separation can be right for the first minute and drift later, when someone moves away from the microphone or a phone line changes quality. Thirty seconds from the middle catches most of it.

If both tracks hold up, you are done: they are yours to download individually and edit anywhere. If something sounds wrong, it will be one of a small number of specific things, and the next lesson is about telling them apart and fixing them with the one control that matters.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Next lessonHow many speakers to declare

All Speaker Separation lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent