A recording of two people in, two clean tracks out. What Speaker Separation costs, what it hands back, and why the first run should leave every setting alone.
Lesson 1 · start · 4 min read
Open Speaker SeparationSpeaker Separation takes one recording with several people in it and gives you back one audio track per person. That is the whole idea, and this lesson runs it once end to end on the easiest case there is: two people, talking in turn.
Note
There is no plan gate on this tool. Every account can run it, including a free one — the only limit is credits. That is worth saying plainly because most of the studio is not like that, and readers assume a lock that is not there.
The two input tabs.
Up to 3 files at a time on the upload tab, each processed as its own job. How large and how long a single file may be depends on your plan, and the tool prints your real figures on the line above the dropzone. Do not go looking for those numbers in a guide: they are read from the server for your account, which is the only place they are ever correct.
There are two controls, and on a first run you want neither of them. The speaker count starts on “Auto detect” and will work out how many people it can hear. The transcript switch starts off. Run it as it is, hear what comes back, and change one thing at a time after that.
Tip
The credit estimate appears beside the input tabs once the length is known. 330 credits per minute of audio you put in — 3,300 credits for ten minutes, 19,800 for an hour, however many people are in it.
The output of a finished job.
Careful
Each track has its own play and download control. There is no zip and no download-all — take the speakers you need individually. If you need all of them, that is several clicks rather than one, and it is worth knowing before you plan a workflow around it.
Two people should give you two. A third track on a two-person recording means one person was heard as two, which the next lesson covers.
You are listening for one voice and one voice only. A second voice under the first is the failure that matters most, and it is the one nothing on screen will tell you about.
A separation can be right for the first minute and drift later, when someone moves away from the microphone or a phone line changes quality. Thirty seconds from the middle catches most of it.
If both tracks hold up, you are done: they are yours to download individually and edit anywhere. If something sounds wrong, it will be one of a small number of specific things, and the next lesson is about telling them apart and fixing them with the one control that matters.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.