🎧 Listen to this article
On This Page
0%- Where AI-narrated audiobooks actually go in 2026
- What you need before you start
- Step 1: Prep the manuscript
- Step 2: Generate the narration
- Step 3: Listen and fix mispronunciations
- Step 4: Master to ACX spec
- Step 5: Produce the opening, closing, and sample
- Step 6: Pick your distribution path and submit
- Step 7: Wait, then go live
- How does this compare to hiring a human narrator?
- FAQ
- What to do next
You finished the manuscript. Now you need an audiobook — and the math on a human narrator (typically $200–$400 per finished hour, plus weeks of pickups) does not pencil out for most indie titles. One thing to get right up front: standard ACX submissions still require human narration, and unauthorized AI or text-to-speech is prohibited. That does not mean AI audiobooks are stuck — it means you take a different path. For Audible, Amazon's own KDP Virtual Voice program lets you create an AI-narrated audiobook; for the rest of the market, you self-upload with disclosure to Apple Books, Google Play Books, and Spotify. By 2026 the production workflow is mature enough that a careful author can produce a retail-quality audiobook in a weekend.
This tutorial walks through the full pipeline: prepping your manuscript, generating narration, mastering to ACX's technical spec, and submitting without triggering a rejection. Every step has an expected output and a fix for the most common failure mode.
TL;DR — the 7 steps:
- Clean your manuscript — split into per-chapter text files, strip footnotes and page numbers.
- Generate the narration in AudioPod's Audiobook Studio, one chapter at a time.
- Listen and re-roll mispronunciations using inline IPA between slashes.
- Master each file to ACX spec: -23 to -18 dB RMS, -3 dB peak, -60 dB noise floor.
- Add opening/closing credits and the retail sample (1–5 minutes from a representative chapter).
- Pick your distribution path — KDP Virtual Voice for Audible, or self-upload with AI disclosure to Apple Books, Google Play Books, and Spotify.
- Submit and wait for review, then go live on your chosen stores.
Where AI-narrated audiobooks actually go in 2026
The most common mistake is assuming ACX opened its standard marketplace to AI narration. It did not. ACX's rule still reads that your submitted audiobook must be narrated by a human unless otherwise authorized — unauthorized AI or third-party text-to-speech is prohibited. The two authorized AI paths within Amazon's world are KDP Virtual Voice (Amazon's own AI-narration program, created through KDP and distributed on Audible — check current eligibility and terms at kdp.amazon.com) and ACX's invite-only narrator Voice Replica beta (for professional narrators replicating their own voices).
Everywhere else, you produce the master file yourself and upload it with a disclosure that the narration is AI-generated. Apple Books, Google Play Books, and Spotify for Authors all accept self-produced AI-narrated audiobooks this way.
The practical effect: a self-published author who used to face a $3,000–$8,000 narration bill for a 10-hour novel can now produce the same audiobook for the cost of an AudioPod Pro subscription and a weekend of editing time — then route it through KDP Virtual Voice for Audible and self-upload to the other stores.
What you need before you start
- The finished manuscript as a
.docxor.txtfile - A free or paid AudioPod account (sign up here — the free tier is enough to test, Pro is enough to finish a typical novel)
- An ACX account (acx.com) with a published Kindle title, or rights to the audio for an existing print/ebook
- A quiet hour to review the master files before submission
You do not need a microphone, an interface, or any DAW experience. The mastering is handled inside AudioPod's DAW and you can export ACX-spec .mp3 files directly.
Step 1: Prep the manuscript
ACX wants one audio file per chapter, plus separate files for the opening credits, closing credits, and retail sample. So your first job is to split the manuscript the same way.
Create a folder structure like this:
12345678
/my-novel/
00-opening-credits.txt
01-chapter-1.txt
02-chapter-2.txt
...
18-closing-credits.txt
19-retail-sample.txt
Strip these from the chapter files before pasting them into the narration tool:
- Page numbers
- Footnotes and endnote markers
- Image captions (unless you want them read aloud)
- Table-of-contents entries
- Copyright page (this goes in the opening credits instead)
Expected output: one .txt file per chapter, each between 1,500 and 8,000 words. If a chapter is longer than 8,000 words, split it at a natural scene break — you can rejoin the audio later in the DAW.
Common error: leaving in [Figure 1.2] style placeholders. The narrator will read them aloud. Search-and-replace these to nothing before generation.
Step 2: Generate the narration
Open Audiobook Studio and create a new project. Pick a voice from the library — for fiction, pick one with a tested narration range; for non-fiction, pick a clearer, more instructional voice. The library shows sample reads, so audition three or four before committing to a 10-hour project.
Paste in chapter 1, set the pace (most fiction reads well at the default), and generate. A 5,000-word chapter takes about 90 seconds to produce.
A few generation settings that matter:
- Pacing: default for most fiction; slower for non-fiction or technical content.
- Pause length: the default 400ms between paragraphs is right for novels. Memoir reads better at 600ms.
- Voice consistency: lock the voice for the whole book before you generate chapter 1 — switching mid-book causes audible tonal shifts even with the same character.
Expected output: a .wav or .mp3 file for each chapter, in the project's outputs panel.
Step 3: Listen and fix mispronunciations
This is the step most authors skip and then get rejected for. Listen to every chapter at 1x speed. Take notes on:
- Proper nouns the narrator mispronounced (character names, invented place names, technical jargon)
- Homographs read with the wrong sense (e.g. "wound" as injury vs past tense of wind)
- Numbers that should be read as words vs digits ("2026" vs "twenty twenty-six")
For each error, fix the pronunciation in AudioPod's editor by typing the word's IPA between slashes, right where the word would go. Instead of writing "Chiron," write its phonetic spelling in place:
12
/ˈkaɪrən/
The engine reads the slash-delimited IPA as the pronunciation and speaks it exactly as written. Then regenerate just the affected paragraph and patch it into the chapter file using the DAW splice tool.
Expected output: every proper noun and homograph reads correctly. Make a per-book pronunciation glossary as you go — for a series, you'll reuse it across every volume.
Common error: trusting the first pass for character names. If your protagonist is named Siân or Xochitl, the narrator will get it wrong. Always re-listen.
Step 4: Master to ACX spec
ACX has a strict technical spec. Files that fail QC are returned and you have to resubmit, which costs you 10–14 days.
| ACX requirement | Target value | How to verify |
|---|---|---|
| Format | MP3, 192 kbps CBR, 44.1 kHz | Export settings in AudioPod |
| Channel | Mono or stereo (mono is fine) | Export settings |
| RMS level | -23 dB to -18 dB | Loudness meter in Audio Reader preview |
| Peak level | ≤ -3 dB | Same meter |
| Noise floor | ≤ -60 dB | Noise Reduction report |
| Room tone | 0.5–1 sec at the start and end of each file | Manual trim in the DAW |
| File length | Each file ≤ 120 minutes | Chapter splitter handles this |
| Opening credits | Title, author, narrator | Generate as a separate file |
| Closing credits | "End of book" + production credits | Generate as a separate file |
In AudioPod, apply these in order on each chapter:
- Run Noise Reduction to push the noise floor below -60 dB. AI-generated narration is usually clean already, but the master pass normalizes any residual hiss.
- Apply the ACX preset in the DAW — this normalizes RMS to -20 dB and caps peaks at -3 dB.
- Add 0.7 sec of room tone before the first word and after the last.
- Export as MP3 192 kbps CBR.
Expected output: every chapter passes the ACX Audio Checker (a free tool from ACX itself — drop your file in and it tells you pass/fail).
Common error: exporting WAV. ACX only accepts MP3. Yes, that loses fidelity. No, you cannot negotiate.
Step 5: Produce the opening, closing, and sample
ACX requires three additional files beyond the chapters:
Opening credits (~15 seconds):
"[Book title], by [Author name]. Narrated by [your chosen voice name or pseudonym]."
Closing credits (~15 seconds):
"This has been [Book title], by [Author name]. Narrated by [voice name]. Thanks for listening."
Retail sample (1–5 minutes): pick the most representative section of the book. For fiction, an early dialogue-heavy scene works best — listeners want to know what the voice sounds like in action, not just declaiming exposition. Splice it together from the existing chapter files in the DAW; don't regenerate, the listener should hear exactly what the audiobook sounds like.
For the narrator name on credits, you can use the voice's display name as listed in Audiobook Studio, or invent a stage name — and always keep your AI-narration disclosure accurate on whichever platform you upload to.
Step 6: Pick your distribution path and submit
Your files are ready. Where they go depends on the store — and this is the step authors most often get wrong.
For Audible → use KDP Virtual Voice, not standard ACX. Standard ACX submissions require human narration, so an indie AI-narrated title goes through Amazon's KDP Virtual Voice program instead: you create the audiobook from your ebook inside KDP, using Amazon's own AI narration, and it distributes to Audible. Availability, eligibility, and royalty terms change over time — confirm the current details at kdp.amazon.com before you plan around it.
For Apple Books, Google Play Books, and Spotify for Authors → self-upload with disclosure. Each store lets you upload your own master files (or go through an approved aggregator) and requires you to disclose that the narration is AI-generated:
- Create or claim your title on the store (or in your aggregator).
- Upload each chapter file in order, plus opening/closing credits and the retail sample where required.
- Set the AI-narration disclosure flag the platform asks for.
- Submit for review.
Expected output: the store confirms your upload and gives you its review window (typically a few business days to two weeks).
Common error: skipping the disclosure. Every platform that accepts AI narration requires it, and a missing or false disclosure can get the title pulled. It does not reduce your royalty rate or hide the book from search.
Step 7: Wait, then go live
Review windows vary by platform — a few business days on some stores, up to ~2 weeks on others. Reviewers check the technical spec, listen to a sample of chapters, and verify the disclosure. If anything fails, you get a specific note and time to fix it, then the title goes live on your chosen stores. Royalty splits and exclusivity terms are set by each platform — check the current rate card before you commit to exclusivity.
How does this compare to hiring a human narrator?
| Path | Cost (10-hr novel) | Turnaround | Pickups | Quality control |
|---|---|---|---|---|
| Human narrator via ACX casting | $2,000–$4,000 PFH royalty share or upfront | 6–12 weeks | 1–3 rounds, $50–$100 each | Narrator handles |
| AI narration via AudioPod Creator | $20/mo subscription | 1–3 days | Self-service, free | You handle |
| AI narration via free-audiobook-generator | Free tier | A few hours for short books | Self-service | You handle |
The human-narrator path still produces a better result for character-heavy literary fiction where the narrator is essentially a co-author. For non-fiction, genre fiction, and series fiction where the priority is shipping the next book, the AI path wins on every axis except prestige.
FAQ
Do I have to disclose that the audiobook is AI-narrated? Yes. Every platform that accepts AI narration — Apple Books, Google Play Books, Spotify for Authors, and Amazon's KDP Virtual Voice — requires an AI-narration disclosure. It does not reduce royalties or affect distribution.
Will Audible reject my book for being AI-narrated? Standard ACX rejects AI narration — its rule requires a human narrator unless otherwise authorized. The authorized route to Audible for an AI-narrated indie title is Amazon's KDP Virtual Voice program (created in KDP, not ACX). Once you're on the correct path, the most common issue is technical (file format, levels, missing room tone) — see step 4.
Can I use AI narration on Spotify Audiobooks and Google Play Books? Yes, both platforms accept AI narration with disclosure. Spotify Audiobooks crossed 200,000 titles in 2026 and a meaningful share are AI-narrated. Google Play Books has been accepting AI narration since 2023 through its auto-narrated program.
What about Kindle Vella and KDP audiobook beta? KDP's audiobook beta (announced for select authors in late 2024) accepts AI narration. Kindle Vella does not require audio. Both pathways are open to indie authors using the workflow above.
Can I narrate a book that's already been narrated by a human? If you hold the audio rights, yes. Many indie authors are re-narrating older titles in AI as a low-cost catalog refresh.
What about multilingual editions? Generate a separate audiobook per language. ACX accepts non-English titles in major languages — check the current list at acx.com. For translation, do that before generation; the narrator only reads what you paste in.
Can I voice-clone myself and narrate as the author? Yes — Creator and above include unlimited custom voice models. Record 5–10 minutes of clean source, train, and use your own voice as the narrator. Disclose AI narration regardless of whose voice you cloned.
What to do next
Start with one chapter. Generate it, master it, run it through the ACX Audio Checker, and listen to the result on a phone speaker (not your studio monitors — most listeners are on phones or earbuds in cars). If chapter 1 sounds right, the rest of the book will too.
The free audiobook generator handles short books on the free tier. For a full novel, Pro at $50/mo covers a typical 10-hour audiobook with credit headroom for pickups. Read the full Audiobook Studio guide for advanced tips on multi-character narration, pacing, and series voice locking.

