The one control that decides whether a separation is right. When automatic is enough, when to declare an exact number, and how to recognise the two ways it goes wrong without telling you.
Lesson 2 · core · 7 min read
Open Speaker SeparationThe speaker count is the only real decision this tool asks you to make, and it is genuinely hard, because the two ways it goes wrong both look like success. The job turns green. The tracks appear. The credits are spent. You find out by listening, or you never find out.
“Auto detect” asks the separation to work out how many distinct voices are in the recording. Declaring a number instead tells it, and it will produce that many. Neither is safer in general. The right one depends entirely on how certain you are.
You know the number for certain
People join or leave partway through
A recording you did not make
Two people, one microphone, one room
A crowd, or heavy crosstalk
The menu offers 2 through 10, so 10 people is the ceiling. That is not an arbitrary cap so much as an honest edge: a recording with more voices than that almost always has them overlapping, and overlapping speech is the thing separation is worst at.
Both of these finish as a successful job. There is no warning, no badge, and no refund. The only detector is your own ear, which is why the first lesson asks you to play thirty seconds of every track before doing anything else.
How a separation disappoints you, and what to do about it.
Careful
A re-run is a new job at the full rate. On a five-minute clip that is 1,650 credits and no reason to hesitate; on a two-hour panel it is 39,600, which is the argument for testing your settings on a short excerpt of a long recording before you commit the whole thing.
Tip
Cut a two-minute excerpt from the middle of a long recording and separate that first. 660 credits buys you the answer to “does the count I think is right actually work here”, and the middle is where crosstalk lives — the opening minute is the least representative part of any conversation.
“Include Speaker-Labeled Transcription” sits under the count and starts off. It adds a timestamped transcript with a speaker label on every line, alongside the separated audio. It costs nothing extra. The credit estimate is the same either way.
That matters more here than it looks, because the transcript is the fastest way to audit a separation you are unsure about. Reading who was credited with which line takes a minute; listening to two full tracks takes as long as the recording. If the transcript shows one person's lines split across two labels, you have found a problem the audio would have taken an hour to reveal.
Careful
Turn it on BEFORE the run. The tool offers no way to add a transcript to a job that has already finished, so getting one afterwards means separating the whole recording again and paying for it again. Since it costs nothing extra, the only reason to leave it off is that you are certain you will never want it.
Not every mistake is worth a re-run. If the tracks are broadly right and a handful of lines landed on the wrong person, fix it in the transcript instead. Move a segment to a different speaker, or to a new one you add. This is how you fix a line that landed on the wrong person. That corrects the record without spending anything.
Note
What a transcript edit does NOT do is move audio. Reassigning a line changes who the transcript says spoke it; the audio for that moment stays on the track the separation put it on. When the audio itself is wrong, a re-run with a declared count is the only fix.
Once the tracks are right, the question becomes what to do with them — which is the next lesson, and the reason most people came here in the first place.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.