Holding one character steady across a long read, writing a script with several voices in it, and knowing the point where Audiobook Studio is the right tool instead.
Lesson 4 · advanced · 8 min read
Open Voice StudioEverything in the first three lessons works on a line. This one is about what breaks when there are four hundred of them — and what breaks is almost never the individual reads. It is consistency: the same character sounding like two people, a passage that stayed dramatic long after the drama ended, a name pronounced two ways in the same script.
Consistency in a long read comes from having FEWER decisions in the script, not more. Every per-line direction is a thing that can disagree with the line above it; a single direction at the top of a passage cannot.
Careful
The carry-forward rule is the one that bites at length. A direction stays in force until another replaces it, so an emotional passage keeps colouring everything after it until you explicitly hand the next segment a plain direction. In a short script you notice within seconds; in a long one you find it on the third listen.
A multi-speaker script is segments with different voices assigned to them. The studio adds a gap between speakers, and that gap is the one setting that exists specifically for this case: 0 to 3 seconds, in steps of 0.1. Default 0.5 seconds.
The controls that matter in a multi-voice script.
Pick a voice per character first and keep it. Casting as you go is how a character ends up with two voices, and the fix is regenerating everything they said.
Two voices of a similar pitch and pace are hard to tell apart in a scene even when they are obviously different in a preview. Audition them back to back, not one at a time.
Too short and the exchange sounds like one person interrupting themselves; too long and the scene loses its pulse. A default gap is right more often than a tuned one.
Give each character their standing direction where their first line appears. It carries forward, so a character with a consistent manner needs it stated once.
The boundary is not a word count. It is whether you are performing a SCRIPT or producing a MANUSCRIPT — whether you will hear every second of the result yourself, or need the tool to tell you which parts are wrong.
Which tool the work belongs in.
Audiobook Studio is the same speech engine with a production workflow around it: it splits a manuscript into chapters and paragraphs, tracks which of them are narrated, listens back and flags the ones that swallowed a word, holds a cast across an entire book, and packages the result to the spec a retailer checks. None of that exists here, and none of it is worth having for a thirty-second script. The Audiobook track covers it end to end.
Tip
If you are unsure, the deciding question is whether you will listen to the whole thing. If yes, stay here — you are the quality check, and this tool is faster. If no, you need a tool that does the checking, and that is the other one.
Long-form synthetic speech has failure modes no setting removes, and it is worth knowing them before you commit an afternoon.
Questions
Free to start
Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.
1,000 credits every month. No card required.