AudioPod AI
  • Pricing

Speaker Separation

Working with the separated tracks

What one track per person makes possible that one mixed track never did, how the labelled transcript earns its place, and the fact that you take them away one at a time.

Lesson 3 · core · 7 min read

Open Speaker Separation

A conversation recorded on one microphone is a single track, and a single track is the reason podcast editing is hard: you cannot touch one person without touching everybody. Separation is worth running because it removes that constraint, and everything in this lesson follows from it.

What separated tracks let you do

Job
What separation changes
Podcast editing
A conversation recorded on one microphone arrives as one track you cannot touch without touching everyone. Separated, each voice can be trimmed, ducked or dropped on its own.
Per-speaker levelling
One person too quiet for the whole recording is unfixable in a mix and trivial on their own track. This is the most common reason to run the tool on audio you were otherwise happy with.
Interview cleanup
Your questions on one track and their answers on another, so you can cut your own interjections out of a long answer without cutting into the answer.
Quoting and clipping
A clean track of one person is what you pull a quote from. With the transcript on, you can find the line first and cut to its timestamp.

Why people run this tool.

Tip

Per-speaker levelling is the one that surprises people. One person too quiet for the whole recording is unfixable in a mix and trivial on their own track. This is the most common reason to run the tool on audio you were otherwise happy with.

What the tool does not do is put them back together. There is no mixer here, no fader, no combined render — separation is where this ends and your editor begins. That is the honest shape of it, and knowing it up front is better than looking for a button that does not exist.

Taking them away

Each speaker has their own player and their own download. Each track has its own play and download control. There is no zip and no download-all — take the speakers you need individually.

Careful

Deleting a separation removes the input file and every separated track, permanently. There is no undo. Download whatever you want to keep before you clear out your history.

The labelled transcript

If you turned the transcript on before the run, it reads in the page under the tracks: every line with a timestamp and the speaker it belongs to, filterable by person and searchable by text. The transcript view shows total speaking time next to each speaker's filter chip, so who talked most is already counted for you.

Its first job is navigation. Finding the sentence you want by reading is faster than scrubbing a waveform, and the timestamp beside it takes you straight to the audio. Its second job is proof: reading who was credited with which line is the quickest audit of whether AudioDiarize put the right voice on the right track.

Correcting it

Edit
What it is for
Reassign a line
Move a segment to a different speaker, or to a new one you add. This is how you fix a line that landed on the wrong person.
Rename a speaker
Type a real name in place of Speaker 1 and it is shown verbatim from then on, in the filter chips and on every line.
Edit the words
Correct a misheard word in place. The timings stay where they are.
Merge or split a segment
Join a segment into the one above it, or cut one in two. What you reach for when the boundary between two turns landed in the wrong place.

What you can change, in the page, without spending credits.

Note

Renaming a speaker is the edit worth doing first and the one people skip. A transcript that says who actually spoke is a document somebody else can read; one that says Speaker 1 and Speaker 2 is a document only you can use.

Getting the transcript out

Format
When you want it
TXT
The transcript as formatted text with speaker labels. What you want for reading, quoting, or pasting into a document.
JSON
Every segment with its speaker, start and end time, plus the detected language and total duration. What you want if something else is going to read it.

The two download formats.

Make your corrections first, then download. A file you saved earlier is a snapshot: a rename or a reassignment you make after saving it is not going to appear in it, and re-downloading is cheaper than reconciling two versions later.

That is the whole tool: a recording in, one track per person and an optional record of who said what out, taken away individually and finished somewhere else. Most of the value is in what you do next, which is exactly why the separation itself should be something you stop thinking about after the second lesson.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonHow many speakers to declare

All Speaker Separation lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent