AudioPod AI
  • Pricing

Audiobook Studio

Narration quality and review

Direct the performance, keep a long book consistent, and prove it is right before you export — including the check that refuses an export when words are missing.

Lesson 3 · core · 9 min read

Open Audiobook Studio

A book is long enough that consistency beats perfection. The goal of a review pass is not a flawless take of every paragraph — it is a performance that holds the same character for eight hours, with nothing in it that a listener would stop and rewind.

Set the performance top-down

Everything in the studio applies in layers: the book-level settings underneath, per-line settings on top. Start at the bottom of that stack, because one change there moves every line and costs you nothing to try.

Control
Range
Default voice
200+ voices, plus any voice you have cloned yourself
Speed
0.8× to 1.2×, in steps of 0.05
Style
Narrative, Conversational, Warm, Dramatic, Calm, Energetic
Narration brief
Free text, up to 600 characters. Leave it empty and the studio writes one from your manuscript.
Voice
Any voice in the picker
Emotion
neutral, warm, tense, sad, excited, intimate, dramatic, authoritative
Voice direction
Free text, e.g. “Low, urgent whisper”
Pause
Auto, 0 ms, 150 ms, 250 ms, 500 ms, 1 s, 2 s

The controls that shape a performance, and where each one lives.

Tip

One paragraph describing how the whole book should be performed. It steers the per-line direction the studio suggests, so it is the cheapest way to move every line at once. Free text, up to 600 characters. Leave it empty and the studio writes one from your manuscript.

The style presets are Narrative, Conversational, Warm, Dramatic, Calm and Energetic, and speed runs 0.8× to 1.2×, in steps of 0.05. Both are project-wide, and both take effect on the next render rather than on audio you already have.

Direct the lines that need directing

Once the book-level performance is right, most paragraphs need nothing. Reach for a per-line control where the text is genuinely ambiguous — a line that is sarcastic on the page, a whispered aside, a shout — and leave the rest alone.

  1. 01

    Voice

    Overrides the narrator for this paragraph only — the control behind a second character reading their own dialogue. Options: Any voice in the picker.

  2. 02

    Emotion

    The feeling behind this line. One choice, applied to the whole paragraph. Options: neutral, warm, tense, sad, excited, intimate, dramatic, authoritative.

  3. 03

    Voice direction

    A short performance note for this line, in your own words — the audiobook equivalent of a margin note to a narrator. Options: Free text, e.g. “Low, urgent whisper”.

  4. 04

    Pause

    How much silence follows this paragraph. Auto lets the engine shorten the gap between lines from the same speaker; the rest are fixed holds. Options: Auto, 0 ms, 150 ms, 250 ms, 500 ms, 1 s, 2 s.

Note

Every per-line control is an instruction for the NEXT render. Changing one on a paragraph that is already narrated does not alter the audio you have — and when it is the direction you changed, the studio badges the row so the mismatch is visible rather than something you discover in the export.

Casting a book with dialogue

Action
What it does
Analyze script
Screenplay Mode. Splits your paragraphs into narration and character lines and writes a per-line direction for each, so every character can carry their own voice.
Auto-detect characters
Finds who speaks in the book and how many lines each of them has, without casting them.
Auto-cast voices
Suggests a distinct voice for every detected character. You can change any of them.

The three casting actions, all in the Cast sheet.

Cast in that order. Analyze script is what turns a wall of narration into narrator lines and character lines; without it there is nothing for the other two to work on. A character called by more than one name — a nickname, an initial, a title — can be pointed at a single canonical voice, so the same person does not arrive in three different voices.

The review pass

Proof by listening, with the text in front of you. The player follows the paragraph rows, so you can hear a line and mark it without losing your place. Flag first and fix in a second pass — stopping to fix each line as you hear it turns a chapter into an afternoon.

  • Pronunciation fix — your own flag, for your own second pass.
  • Bad emotion — your own flag, for your own second pass.
  • Wrong voice — your own flag, for your own second pass.
  • Regenerate later — your own flag, for your own second pass.

Two badges appear without you asking for them, written by the studio rather than by you. They are the ones worth understanding, because one of them can stop an export.

Badge
What it means
May be missing words
The studio listened back to the take and compared it against your manuscript. This badge means some words were not audible. An export refuses while these are unresolved, so a book cannot ship with a paragraph that swallowed a sentence.
Direction changed after narration
You changed the emotion or direction after the audio was rendered, so what you are hearing no longer matches what the row says. Re-narrate that line to catch it up.

Badges the studio puts on a row by itself.

Careful

“May be missing words” is a hard gate, not a suggestion: the export refuses while any paragraph is flagged, until you re-narrate it or explicitly export anyway. It exists because a missing sentence is the one defect a listener always notices and an author almost never does.

Prove it before you export

The export sheet runs a readiness checklist first. It separates what a retailer will reject from what is merely optional, so you are not guessing which warning matters.

Check
Blocks submission?
Every chapter narrated
Yes — fix before you submit
Book title set, and not just the filename
Yes — fix before you submit
Cover art at 2400×2400
Yes — fix before you submit
Author name
No — advisory
Spoken opening and closing credits
No — advisory
Retail sample source and length
No — advisory
Loudness normalized — handled during the export
No — advisory

The readiness checklist, and what actually blocks a submission.

Then choose what to package. MP3 — ACX-spec (192 kbps) and WAV — lossless master — the first is the submission package, the second is a lossless master for archiving or for editing elsewhere.

Extra
What it adds
Chaptered M4B
A single file with chapter markers, alongside the per-chapter MP3s. Handy for sending the whole book to a reviewer.
Retail sample
A short excerpt cut from a chapter you choose — 1 minute, 2 minutes, 3 minutes, 5 minutes.
AI voice disclosure
Records in the package that the narration was synthesized. On by default, and worth leaving on: retailers increasingly ask.
Cover art
A square JPG or PNG, at least 2400×2400, included in the package as cover.jpg.

The optional extras on a submission package.

Room tone is handled for you. Every exported chapter is padded with 0.9 seconds of silence at the head and 3 at the tail — inside the 0.5–1 second opening and 1–5 second closing window a retailer expects — and you can change both if you have a reason to. Loudness is normalized during the export, so that is one spec you never have to think about.

Tip

Cut the retail sample from a chapter that shows the book at its best rather than from chapter one out of habit. The lengths on offer are 1 minute, 2 minutes, 3 minutes and 5 minutes, and it is the only part of the book most browsers will ever hear.

A book that clears the checklist, has no outstanding “May be missing words” flags and sounds like one narrator throughout is a book you can submit. The cover is the item people leave to last and the one a retailer rejects first — no amount of good narration substitutes for it.

Questions

What people ask about this

Previous lessonPreparing your manuscript

All Audiobook Studio lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent