🎧 Listen to this article
On This Page
0%- TL;DR — the five things ACX checks
- Step 1 — Prep the manuscript for narration
- Step 2 — Choose and lock a narrator voice
- Step 3 — Render chapter by chapter
- Step 4 — Master to ACX loudness spec
- Step 5 — Build the required structure files
- Step 6 — Declare AI narration and upload
- What model does AudioPod use to narrate?
- Common errors and fixes
- How does this compare to hiring a narrator or other tools?
- FAQ
- Start your first chapter free
You wrote the book. Now ACX is rejecting your upload for a noise floor that's 4 dB too loud, or a chapter file that runs 3 minutes over the limit — and the rejection email won't tell you which file failed. The technical bar is the part that trips up most first-time self-publishers, not the narration itself.
This is a step-by-step walkthrough to get an AI-narrated audiobook through ACX's audio spec and disclosure rules on the first submission. It assumes you have a finished manuscript and an Audiobook Studio project ready to render.
TL;DR — the five things ACX checks
- Format: 192 kbps or higher, constant bitrate MP3, 44.1 kHz — one file per chapter.
- Loudness: RMS between −23 dB and −18 dB, peaks no higher than −3 dB.
- Noise floor: −60 dB RMS or quieter (this is where AI narration wins automatically).
- Structure: opening credits, closing credits, and a standalone retail sample.
- Disclosure: as of 2026, ACX accepts AI narration if you declare it during setup.
Get those five right and the rest is metadata. Below is the full process.
Step 1 — Prep the manuscript for narration
Clean text produces clean audio. Before you render anything:
- Strip headers, footers, and page numbers — they read aloud as garbage.
- Spell out anything you want pronounced a specific way ("Dr." → "Doctor", "1996" → "nineteen ninety-six" if the read matters).
- Mark chapter breaks clearly. Each ACX chapter becomes its own file, so your chapter structure is your file structure.
- Decide on opening and closing credits text now. ACX requires both. Opening: title, author, narrated by. Closing: "This has been [title], written by [author], narrated by [narrator]. Production copyright [year]."
Expected output: a chapter-segmented manuscript where every break maps to a file you'll export.
Step 2 — Choose and lock a narrator voice
Consistency across hours of audio is the hard part of narration, and it's where AI has a structural edge: the voice never drifts, never has an off day, and never needs a pickup session three weeks later. Browse voices in the AI narrator library and audition two or three against a paragraph of your actual prose — dialogue-heavy fiction stresses a voice differently than steady nonfiction.
Lock one voice for the whole book. Switching voices mid-title is the fastest way to get a quality rejection from ACX's human reviewers.
If you want to hear the difference before committing, the free audiobook generator lets you render a sample chapter on the free tier — no card required.
Step 3 — Render chapter by chapter
Render each chapter as a separate job. Two reasons:
- ACX wants one file per chapter anyway.
- If chapter 7 has a pronunciation you want to fix, you re-render 7 — not the whole book.
In Audiobook Studio, generate each chapter, then listen end-to-end at 1.5× speed. You're listening for the things automated checks miss: a number read as digits when you wanted words, an acronym spelled out wrong, a homograph ("read", "lead", "bass") that landed on the wrong pronunciation. Fix those in the text and re-render the single chapter.
Expected output: one clean audio file per chapter, plus a separate opening-credits file and closing-credits file.
Step 4 — Master to ACX loudness spec
This is where most rejections happen. ACX's published technical requirements (current as of 2026 — always confirm against ACX's own help pages before you submit):
| Requirement | ACX spec | Why it fails |
|---|---|---|
| Peak level | ≤ −3 dB | Over-loud master clips |
| RMS (average loudness) | −23 dB to −18 dB | Too quiet or too hot |
| Noise floor | ≤ −60 dB RMS | Room hiss on human-recorded tracks |
| Format | MP3, 192 kbps CBR, 44.1 kHz | Variable bitrate or low kbps |
| File length | ≤ 120 minutes each | Long chapters need splitting |
The noise-floor requirement is the one that eats human narrators alive — a quiet room is rarely quiet enough. AI-rendered narration has effectively no room tone, so it clears −60 dB without any cleanup. If a chapter still reads hot or quiet on RMS, normalize it: target around −20 dB RMS to sit safely in the middle of the band, and confirm peaks stay under −3 dB.
If you recorded any human pickups and they carry hiss, run them through noise reduction before mastering so the whole title matches.
Step 5 — Build the required structure files
ACX wants more than chapters. Assemble:
- Opening credits — short standalone file.
- Closing credits — short standalone file.
- Retail sample — 1 to 5 minutes pulled from the body of the book (not the intro). Pick a passage that sells the read.
Render these from the same locked voice so they match the book.
Step 6 — Declare AI narration and upload
During ACX project setup, you'll choose how the title was produced. As of 2026, ACX and Audible accept AI-narrated audiobooks with disclosure — you declare that the narration was AI-generated during setup. Do not skip this. Misrepresenting an AI narration as a human read is the kind of thing that gets a title pulled after it's live.
Upload your files in order: opening credits, chapters 1–N, closing credits, then the retail sample in its own slot.
Expected output: a project that passes ACX's automated audio check and moves into review.
What model does AudioPod use to narrate?
We use AudioPod's proprietary audio AI stack — we don't share specific vendor or model details. What matters for ACX is the output: clean, consistent narration with a noise floor well under the −60 dB requirement and stable loudness across every chapter, which is exactly what the spec above asks for.
Common errors and fixes
| Rejection / problem | Cause | Fix |
|---|---|---|
| "File exceeds maximum length" | Chapter over 120 min | Split at a natural scene break into two files |
| "Audio is too loud / peaks over −3 dB" | Over-normalized master | Re-normalize to −20 dB RMS, cap peaks at −3 dB |
| "Noise floor too high" | Human pickup with room tone | Run the pickup through noise reduction |
| Mispronounced name or number | Ambiguous source text | Spell it phonetically in the manuscript, re-render that chapter |
| Inconsistent voice between chapters | Voice changed mid-project | Re-render odd chapters with the locked voice |
| Bitrate rejected | Exported as VBR or under 192 kbps | Re-export as 192 kbps CBR MP3 |
How does this compare to hiring a narrator or other tools?
Human narration runs roughly $200–$400 per finished hour and weeks of scheduling, per common industry rates. AI narration collapses that to minutes per chapter and makes pickups free — you just re-render. Among AI tools, ElevenLabs, NarrationBox, and Speechify all produce long-form narration; the differences that matter for ACX are per-project cost, whether mastering is built in, and how long your files stay retrievable for re-edits.
AudioPod starts free — 1,000 credits a month, every tool, no card — with paid plans from $20/mo (Creator). See pricing for the full breakdown. Generated files are retained on a tier-based ladder (Free 1 year, Creator 2 years, Pro 3 years), and your job history stays in the dashboard so you can re-render any chapter in one click even after the audio ages out — useful when ACX asks for a correction months after release.
FAQ
Does ACX really accept AI-narrated audiobooks in 2026? Yes, with disclosure. You declare AI narration during project setup. Misrepresenting it as human is the violation, not the AI itself.
Why do AI narrations pass the noise-floor check so easily? There's no microphone, no room, no HVAC hum. The audio is generated clean, so it sits well under the −60 dB requirement that defeats most home studios.
Can I mix human and AI narration in one title? Technically possible, but match loudness and noise floor across both or ACX's reviewers will flag the inconsistency. Run human segments through noise reduction first.
How long should the retail sample be? One to five minutes, pulled from the body of the book — not the credits or intro. Choose a passage that demonstrates the narration at its best.
What's the fastest way to fix a single mispronounced word? Edit the source text for that chapter (spell the word phonetically), then re-render only that chapter. You don't touch the rest of the book.
Do I need separate files for credits? Yes. Opening credits, closing credits, and the retail sample are each their own upload slot, separate from your chapter files.
Start your first chapter free
The technical bar is real, but it's a checklist, not a craft. Render a clean voice, master to −20 dB RMS with peaks under −3 dB, structure your files, disclose the AI narration, and you'll clear ACX on the first pass.
Try it on the free audiobook generator — render a sample chapter on the free tier, listen to the noise floor for yourself, and decide before you spend a credit. When you're ready for the full book, Audiobook Studio handles every chapter end to end.

