The three settings that decide whether a transcript is usable: when premium accuracy is worth paying for, how to get speaker labels you can trust, and what word timestamps unlock.
Lesson 2 · core · 7 min read
Open TranscriptionThree settings do almost all the work here. One decides how well the words are heard, one decides whether the transcript knows who is talking, and one decides whether it knows exactly when. Everything else in the panel is a detail by comparison.
AudioTranscribe offers two accuracy tiers, and the honest way to choose is to judge the RECORDING rather than the importance of the job. A clear voice on a decent microphone comes back much the same either way; a noisy four-way call does not.
The two tiers, and who can select them.
One person, quiet room, headset mic
Interview, two mics, some crosstalk
Conference call recorded through a laptop
A transcript that will be read, not published
Note
Premium is a plan feature: Available on Creator and above, from $20/mo. On a free account the option is visible with a lock rather than hidden, so you can see what you are choosing between — and Standard is available to everyone, on every plan.
Splits the transcript by who is talking and labels each segment with a speaker. On by default. Turn it off for a recording with one voice — there is nothing to separate, and a single speaker occasionally gets split in two. It is the setting people notice most, because a transcript that labels the wrong person is more annoying than a transcript that misspells a word.
Declares how many people are in the recording instead of letting it work that out. Only appears while diarization is on. Auto-detect, or an exact count from the picker — and if you know the number, say it. Separation is a judgement about how many distinct voices are present, and removing that judgement removes the two failure modes that come with it.
Tip
Turn separation OFF for a single-voice recording. There is nothing to split, and the only thing separation can do on a solo dictation is invent a second speaker.
Careful
Do not declare a count you are guessing at. Auto-detect being wrong is recoverable by editing a few segments; a confidently wrong count reshapes the whole transcript around a number that was never true.
Records a start and end time for every word rather than for each segment. This is what makes tight captions and word-level search possible later. They change nothing about the words you read on screen, which is why they are easy to dismiss — the difference appears in what you can do with the transcript afterwards.
The next lesson is about what happens after all this: fixing what is left by hand, and choosing the export format that suits where the transcript is going.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.