AudioPod AI
  • Pricing

Loading blog...

AudioPod AI
  • Pricing

Loading article...

AudioPod AI
  • Pricing
AudioPod AI
  • Pricing

Loading article...

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Noise Reduction guides
  • Transcription guides
  • Voice Changer guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent
Directing markup — an [excited] tag highlighted over a waveform on an indigo gradient
HomeBlogProduct

AI Voice Directing: Control Emotion, Pauses, and Pronunciation in Text to Speech

Most text to speech reads your words. AudioSonic Premium lets you direct them — emotion, non-verbal sounds, timed pauses, and exact pronunciation, all with simple inline markup. Hear it in action.

AudioPod Team
•Product•July 9, 2026•4 min read

🎧 Listen to this article

On This Page

0%
  • Direct the emotion
  • Add real, non-verbal sounds
  • Place pauses exactly where you want the beat
  • Fix any pronunciation
  • Follow along, word by word
  • Design or clone the voice itself
  • Direct in 100+ languages
  • How to write directions that actually land
  • Try it

Most text-to-speech engines have one job: read your words out loud. That is fine for a subway announcement. It is not fine for an audiobook, an ad, a character, or a podcast — anything where how a line is delivered matters as much as what it says.

AudioSonic Premium treats a script the way a director treats an actor. You do not just type text; you direct the performance — emotion, pacing, breaths, and exact pronunciation — with a tiny inline markup vocabulary. Every clip below was generated with the real engine. Press play.

Direct the emotion

Put a direction in square brackets at the very start of a line and the voice performs the entire line that way. Here is the same sentence, three ways — only the bracketed direction changed:

[speak excitedly at a fast pace] We actually did it. After all this time, we actually did it.

Excited

Voice Aanya · AudioSonic Premium

[whisper in a hushed awed tone] We actually did it. After all this time, we actually did it.

Whispering

Voice Aanya · AudioSonic Premium

[speak slowly and somberly with a heavy heart] We actually did it. After all this time, we actually did it.

Somber

Voice Aanya · AudioSonic Premium

There is no fixed list of emotions to pick from. You write the direction in natural language, and you can layer dimensions — mood, rhythm, pitch, and vocal style — in a single bracket: [say sadly with deliberate pauses in a low hushed voice].

Add real, non-verbal sounds

Some moments need a laugh, a sigh, or a breath — not a word. Drop a sound tag anywhere in the text and it renders as an actual sound:

That's hilarious [laugh] okay [sigh] let me catch my breath [breathe] and get back to it.

Non-verbal sounds

Voice Aanya · AudioSonic Premium

The insertable sounds are [laugh], [sigh], [clear throat], [breathe], [cough], and [yawn].

Place pauses exactly where you want the beat

Comedic timing, dramatic reveals, and natural narration all live in the pauses. Insert a timed one with a break tag — up to 10 seconds each:

And the winner is <break time="1s"/> well, you already know. <break time="500ms"/> Congratulations.

Timed pauses

Voice Aanya · AudioSonic Premium

Fix any pronunciation

Names, brands, and homographs are where generic TTS embarrasses you. Override the pronunciation by typing the phonetics (IPA) between slashes, in place of the word:

The founder's name is /ˈraːkeɪʃ/, and the product is /ˈɔːdioʊpɒd/. Say them right every time.

Pronunciation control

Voice Aanya · AudioSonic Premium

Follow along, word by word

Turn on word timestamps and every generation comes back with a synced transcript — the current word highlights as it plays. It is the backbone of karaoke-style captions, video subtitles, and read-along experiences, and it lines up to the audio with no manual timing.

Design or clone the voice itself

Directing shapes how a voice performs. You can also choose which voice:

  • Design a voice from a plain description ("a warm, seasoned audiobook narrator") and get instant preview candidates.
  • Clone a voice from a short sample and reuse it across every tool, in any supported language.
[speak warmly like a seasoned audiobook narrator] Once upon a quiet evening, the story began.

Designed: warm narrator

Voice Aanya · AudioSonic Premium

Direct in 100+ languages

The same directing, cloning, and design controls work across more than 100 languages and locales:

Spanish

Voice Aarav · AudioSonic Premium

Japanese

Voice Abby · AudioSonic Premium

How to write directions that actually land

A few rules get you dramatically better results:

  1. Put one direction at the very start of a segment. To change the delivery mid-script, start a new segment with a new direction.
  2. Write it in English, even for non-English text.
  3. Lowercase, no punctuation inside the bracket. [speak softly and slowly] beats [Speak Softly, and Slowly.].
  4. Be descriptive. The more you describe the performance — mood, pace, pitch, vocal style — the more nuanced the result.
  5. Avoid contradictions. Do not ask for a whisper and a shout in the same breath.

Try it

Open the Text to Speech studio and direct your first line — it is free to start, no card required. When you are ready for unlimited custom voices and higher limits, the pricing starts at $20/mo for Creator.

The words are yours. Now the performance is too.

Tags

#text-to-speech#ai-voices#voice-directing#emotional-tts#voice-cloning

Share this article

On This Page

0%
  • Direct the emotion
  • Add real, non-verbal sounds
  • Place pauses exactly where you want the beat
  • Fix any pronunciation
  • Follow along, word by word
  • Design or clone the voice itself
  • Direct in 100+ languages
  • How to write directions that actually land
  • Try it

Related Articles

AudioPod for Agents: One MCP Server, Two SDKs, One CLI
Product
May 6, 20265 min read

AudioPod for Agents: One MCP Server, Two SDKs, One CLI

Today we're shipping AudioPod for Agents — a real Streamable-HTTP MCP server at mcp.audiopod.ai, a bundled CLI in both SDKs, and a single landing page that documents every developer surface AudioPod publishes.

Read article
How to Make an Audiobook Retail Sample That Sells
Tutorials
August 4, 202611 min read

How to Make an Audiobook Retail Sample That Sells

The retail sample gets more plays than the rest of your audiobook combined. A step-by-step workflow for building one that converts browsers into buyers.

Read article
Inside Audiobook Studio: How a Manuscript Becomes a Directed Audiobook
Features
August 4, 20266 min read

Inside Audiobook Studio: How a Manuscript Becomes a Directed Audiobook

A text-to-speech tool reads your book. A studio produces it. This is the full walkthrough of Audiobook Studio: manuscript parsing, AI voice direction that writes performance notes for every paragraph, casting across 200+ voices, verified narration, and ACX-spec masters you own outright.

Read article

Try AudioPod free

Turn this into your own audio — start free, no card required.

Get started freeSee pricing

Get free audio tips and early access

The best of AudioPod in your inbox — no spam, unsubscribe anytime.

Weekly audio tips · Feature previews · Exclusive discounts