What one track per person makes possible that one mixed track never did, how the labelled transcript earns its place, and the fact that you take them away one at a time.
Lesson 3 · core · 7 min read
Open Speaker SeparationA conversation recorded on one microphone is a single track, and a single track is the reason podcast editing is hard: you cannot touch one person without touching everybody. Separation is worth running because it removes that constraint, and everything in this lesson follows from it.
Why people run this tool.
Tip
Per-speaker levelling is the one that surprises people. One person too quiet for the whole recording is unfixable in a mix and trivial on their own track. This is the most common reason to run the tool on audio you were otherwise happy with.
What the tool does not do is put them back together. There is no mixer here, no fader, no combined render — separation is where this ends and your editor begins. That is the honest shape of it, and knowing it up front is better than looking for a button that does not exist.
Each speaker has their own player and their own download. Each track has its own play and download control. There is no zip and no download-all — take the speakers you need individually.
Careful
Deleting a separation removes the input file and every separated track, permanently. There is no undo. Download whatever you want to keep before you clear out your history.
If you turned the transcript on before the run, it reads in the page under the tracks: every line with a timestamp and the speaker it belongs to, filterable by person and searchable by text. The transcript view shows total speaking time next to each speaker's filter chip, so who talked most is already counted for you.
Its first job is navigation. Finding the sentence you want by reading is faster than scrubbing a waveform, and the timestamp beside it takes you straight to the audio. Its second job is proof: reading who was credited with which line is the quickest audit of whether AudioDiarize put the right voice on the right track.
What you can change, in the page, without spending credits.
Note
Renaming a speaker is the edit worth doing first and the one people skip. A transcript that says who actually spoke is a document somebody else can read; one that says Speaker 1 and Speaker 2 is a document only you can use.
The two download formats.
Make your corrections first, then download. A file you saved earlier is a snapshot: a rename or a reassignment you make after saving it is not going to appear in it, and re-downloading is cheaper than reconciling two versions later.
That is the whole tool: a recording in, one track per person and an optional record of who said what out, taken away individually and finished somewhere else. Most of the value is in what you do next, which is exactly why the separation itself should be something you stop thinking about after the second lesson.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.