AudioPod AI
  • Pricing

Voice Studio

Speaking in any voice.

From typed text to a performance you'd keep — directing delivery, cloning a voice with consent, and holding one character across a long read.

Start with lesson oneOpen Voice Studio

The curriculum

Read them in this order

The second lesson is the one that changes what you think this tool is. Cloning is deliberately third — a voice you clone to fix a performance problem is a slot spent on something one bracket would have solved.

  1. 01Your first voicePick a voice, type a line, hear it back. One success in about a minute — and the two things worth noticing while it plays.start · 4 min read
  2. 02Directing a performanceThe markup that turns a correct read into the read you wanted — bracketed directions, sound tags, timed pauses and phonetic spelling, and what each one actually changes.core · 9 min read
  3. 03Cloning a voiceWhose voice you may clone and why that comes first, what makes a reference recording good, how many clones each plan holds, and the cases where a clone is the wrong tool.core · 8 min read
  4. 04Long form and multiple speakersHolding one character steady across a long read, writing a script with several voices in it, and knowing the point where Audiobook Studio is the right tool instead.advanced · 8 min read

What you are steering

The markup, briefly

Six devices, typed into the script itself. The second lesson explains what each one changes and where each one stops — this is the list, so you know what exists.

Directions
A bracketed acting note at the very start of a line. It steers how the line is delivered and is never spoken. One per line, in English, lowercase, no punctuation — and it carries forward to the lines after it until another direction replaces it.
Sound tags
A non-verbal sound, dropped anywhere in a line. It renders as an actual sound rather than as spoken words, so it is the one bracket you can put mid-sentence.
Emphasis
Capitalise a word — or one syllable of it — to stress it. It works only inside the words you want spoken, never inside a direction, and it stops meaning anything if you use it on every other word.
Pauses
An explicit silence of a length you choose. Whole seconds render as seconds and fractions render as milliseconds, so a half-beat is as easy to ask for as a long one.
Pronunciation
Phonetic spelling that replaces a word. You type IPA between slashes INSTEAD OF the word, not next to it — this is the fix for a name, a homograph, or a technical term the voice keeps getting wrong.
Punctuation & pacing
Not markup at all, and the reason most lines that need fixing do not need markup. Full stops and commas set the natural pauses, an em-dash adds a beat, and a paragraph break rests longer than any of them.

Producing a whole book rather than a script? The Audiobook Studio track covers manuscripts, casting across chapters, and packaging to a retailer’s spec.

Free to start

Read it with the studio open

You do not need a plan to work through this track. A free account gets you into the studio — every control in these lessons, cloning included, is open on the free tier.

Create a free account

1,000 credits every month. No card required.

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent