AudioPod AI
  • Pricing

Voice Studio

Directing a performance

The markup that turns a correct read into the read you wanted — bracketed directions, sound tags, timed pauses and phonetic spelling, and what each one actually changes.

Lesson 2 · core · 9 min read

Open Voice Studio

Everything in this lesson is typed into the script itself. There is no separate performance panel — the direction, the pauses and the pronunciation live in the text, next to the words they apply to, which is why they survive copying a script from one project to another and why they are worth learning as writing rather than as settings.

Note

This is the lesson that changes what people think the tool is. A reader who leaves after lesson one has a text-to-speech box. A reader who finishes this one has a performance they directed — and the gap between those two is almost entirely bracket syntax.

The whole vocabulary

There are six devices and they are all you get. That is a feature: a small vocabulary you know completely beats a large one you half-remember, and each of these does exactly one thing.

Device
What it does
Directions
A bracketed acting note at the very start of a line. It steers how the line is delivered and is never spoken. One per line, in English, lowercase, no punctuation — and it carries forward to the lines after it until another direction replaces it.
Sound tags
A non-verbal sound, dropped anywhere in a line. It renders as an actual sound rather than as spoken words, so it is the one bracket you can put mid-sentence.
Emphasis
Capitalise a word — or one syllable of it — to stress it. It works only inside the words you want spoken, never inside a direction, and it stops meaning anything if you use it on every other word.
Pauses
An explicit silence of a length you choose. Whole seconds render as seconds and fractions render as milliseconds, so a half-beat is as easy to ask for as a long one.
Pronunciation
Phonetic spelling that replaces a word. You type IPA between slashes INSTEAD OF the word, not next to it — this is the fix for a name, a homograph, or a technical term the voice keeps getting wrong.
Punctuation & pacing
Not markup at all, and the reason most lines that need fixing do not need markup. Full stops and commas set the natural pauses, an em-dash adds a beat, and a paragraph break rests longer than any of them.

Every markup device, and what it changes.

Directions

A bracketed acting note at the very start of a line. It steers how the line is delivered and is never spoken. One per line, in English, lowercase, no punctuation — and it carries forward to the lines after it until another direction replaces it. It is the highest-leverage thing on this page: one bracket moves the whole read, where every other device adjusts a detail.

A directed line

Recipe · paste as-is

Style description

[say quietly, with a slow pace] We are going to be fine.

Up to 120 characters inside the brackets. Go over and it stops being recognised as a direction — which means it gets read aloud. Write it descriptively — mood, pace and manner in one bracket — and avoid asking for two contradictory things at once, because a direction to read slowly and urgently resolves into neither.

Careful

Directions carry forward. A direction opens a passage and stays in force until a different one replaces it, so the way to end an emotional stretch is an explicit plain direction on the next segment — not leaving the brackets off, which changes nothing.

Sound tags

A non-verbal sound, dropped anywhere in a line. It renders as an actual sound rather than as spoken words, so it is the one bracket you can put mid-sentence. There are 6: [laugh], [sigh], [clear throat], [breathe], [cough] and [yawn]. Place one where a real reader would actually make that sound — after a hard line, before a difficult admission — and it reads as a person. Place one every third sentence and it reads as a machine imitating a person, which is worse than the flat version you started with.

Note

[chuckle], [gasp], [groan] still work if you type them — older scripts contain them — but they are not offered as chips, so they are the least reliable of the set.

Careful

A bracket at the very start of a line is read as a direction — unless what is inside it is one of the 6 sound tags, in which case it stays a sound. So a line opening with [laugh] gets a laugh, not a delivery instruction, and any other opening bracket steers the read instead of being spoken.

Pauses

An explicit silence of a length you choose. Whole seconds render as seconds and fractions render as milliseconds, so a half-beat is as easy to ask for as a long one. Each pause is 0.1–10 seconds, and a single generation accepts 20 of them.

Timed pauses, long and short

Recipe · paste as-is

Style description

Take a breath. <break time="1s"/> Now continue. <break time="500ms"/> Almost.

Reach for these only when punctuation cannot do it. A full stop, a comma and an em-dash already produce natural rests, and a script full of explicit break tags where commas would have done sounds mechanically spaced — the timing is too exact to be human.

Emphasis

Capitalise a word — or one syllable of it — to stress it. It works only inside the words you want spoken, never inside a direction, and it stops meaning anything if you use it on every other word. It works on part of a word as well as a whole one, which is the useful part: capitalising a single syllable stresses that syllable rather than shouting the word.

Stressing a word and a syllable

Recipe · paste as-is

Style description

We NEED this. AbsoLUTEly.

Capitals inside a direction do nothing — the bracket is an instruction, not speech. Emphasis belongs in the words you want spoken.

Pronunciation

Phonetic spelling that replaces a word. You type IPA between slashes INSTEAD OF the word, not next to it — this is the fix for a name, a homograph, or a technical term the voice keeps getting wrong. This is the device that fixes the problem people give up on: a brand name, a surname, or a word that is spelled the same as another word and read as the wrong one.

Word
Phonetic spelling
create
/kriːt/
read (past)
/rɛd/
live (verb)
/lɪv/
lead (metal)
/lɛd/

The classic cases — one mispronounced word, three homographs.

Tip

Fix a pronunciation the first time you hear it wrong. A name read incorrectly in your second segment will be read incorrectly in your fortieth, and it is the same single token either way.

Punctuation & pacing

Not markup at all, and the reason most lines that need fixing do not need markup. Full stops and commas set the natural pauses, an em-dash adds a beat, and a paragraph break rests longer than any of them. Worth stating explicitly because it is the answer to most "how do I make it pause here" questions, and it is free: rewriting the sentence is usually a better fix than instrumenting it.

Write this
Not this
Getting a beat before a reveal
An em-dash. It produces a natural beat and it survives being read by a person.
An explicit break tag between every clause, timed by hand.
Making a line sound urgent
A direction asking for urgency. Speed changes pace only; a direction changes the performance, and the two do not sound alike.
Raising the speed setting until it sounds hurried.
Stressing a point
Capitalising the one word that carries the meaning.
Capitalising most of the sentence.
Fixing a mispronounced surname
Replacing it with its phonetic spelling between slashes, which is unambiguous rather than another guess.
Respelling it phonetically in ordinary letters and hoping.

Getting a beat before a reveal

Write this:
An em-dash. It produces a natural beat and it survives being read by a person.
Not this:
An explicit break tag between every clause, timed by hand.

Making a line sound urgent

Write this:
A direction asking for urgency. Speed changes pace only; a direction changes the performance, and the two do not sound alike.
Not this:
Raising the speed setting until it sounds hurried.

Stressing a point

Write this:
Capitalising the one word that carries the meaning.
Not this:
Capitalising most of the sentence.

Fixing a mispronounced surname

Write this:
Replacing it with its phonetic spelling between slashes, which is unambiguous rather than another guess.
Not this:
Respelling it phonetically in ordinary letters and hoping.

The three numeric controls

Everything above is typed into the script. These three are sliders, and they are worth knowing because two of them are routinely used to attempt something the markup does better.

Control
Range
Speed
0.5× to 2×, in steps of 0.1. Default 1×.
Pitch shift
-12 to +12 semitones, in whole steps. Default 0.
Silence duration
0 to 3 seconds, in steps of 0.1. Default 0.5 seconds.

Bounds as the studio enforces them.

How fast the voice reads. It changes pace without changing pitch, so a faster read still sounds like the same person — which is what makes it safe to nudge and unsafe to lean on. Moves the voice up or down in whole semitones. This belongs to the voice changer, not to reading typed text, and it is a transposition rather than a different performance. The gap left between one speaker and the next in a multi-voice script. It only appears once your script actually contains more than one voice, which is why you may never have seen it.

Careful

Speed is not a performance control. Pushing it to its ceiling makes a neutral read into a fast neutral read, not an excited one — and a listener hears the difference immediately. If you want energy, ask for energy in a direction and leave speed near its default.

What the markup cannot do

Saying where a tool stops is more useful than another example of where it works, and these are the four requests that come back disappointed.

  • You cannot direct part of a line. A direction applies from where it opens, so a sentence that needs to change halfway needs to be two segments.
  • You cannot ask for a specific pitch contour or a musical note. Pitch shift transposes the whole voice; there is no per-word melody control.
  • You cannot stack contradictory directions and get a blend. Slow and fast in one bracket does not average — it resolves into whichever the engine finds first, and the result is nobody's intent.
  • You cannot make a catalogue voice into a different person. Direction changes performance, not identity. That is what cloning is for, and the next lesson is about when it is genuinely the answer.

Tip

Change one thing per regeneration. Four edits at once tells you the line got better without telling you which edit did it — and the one that made no difference is now permanently in your script, being copied into the next one.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonYour first voiceNext lessonCloning a voice

All Voice Studio lessons · Producing a whole book instead

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent