🎧 Listen to this article
Most text-to-speech engines have one job: read your words out loud. That is fine for a subway announcement. It is not fine for an audiobook, an ad, a character, or a podcast — anything where how a line is delivered matters as much as what it says.
AudioSonic Premium treats a script the way a director treats an actor. You do not just type text; you direct the performance — emotion, pacing, breaths, and exact pronunciation — with a tiny inline markup vocabulary. Every clip below was generated with the real engine. Press play.
Direct the emotion
Put a direction in square brackets at the very start of a line and the voice performs the entire line that way. Here is the same sentence, three ways — only the bracketed direction changed:
Excited
Voice Aanya · AudioSonic Premium
Whispering
Voice Aanya · AudioSonic Premium
Somber
Voice Aanya · AudioSonic Premium
There is no fixed list of emotions to pick from. You write the direction in natural language, and you can layer dimensions — mood, rhythm, pitch, and vocal style — in a single bracket: [say sadly with deliberate pauses in a low hushed voice].
Add real, non-verbal sounds
Some moments need a laugh, a sigh, or a breath — not a word. Drop a sound tag anywhere in the text and it renders as an actual sound:
Non-verbal sounds
Voice Aanya · AudioSonic Premium
The insertable sounds are [laugh], [sigh], [clear throat], [breathe], [cough], and [yawn].
Place pauses exactly where you want the beat
Comedic timing, dramatic reveals, and natural narration all live in the pauses. Insert a timed one with a break tag — up to 10 seconds each:
Timed pauses
Voice Aanya · AudioSonic Premium
Fix any pronunciation
Names, brands, and homographs are where generic TTS embarrasses you. Override the pronunciation by typing the phonetics (IPA) between slashes, in place of the word:
Pronunciation control
Voice Aanya · AudioSonic Premium
Follow along, word by word
Turn on word timestamps and every generation comes back with a synced transcript — the current word highlights as it plays. It is the backbone of karaoke-style captions, video subtitles, and read-along experiences, and it lines up to the audio with no manual timing.
Design or clone the voice itself
Directing shapes how a voice performs. You can also choose which voice:
- Design a voice from a plain description ("a warm, seasoned audiobook narrator") and get instant preview candidates.
- Clone a voice from a short sample and reuse it across every tool, in any supported language.
Designed: warm narrator
Voice Aanya · AudioSonic Premium
Direct in 100+ languages
The same directing, cloning, and design controls work across more than 100 languages and locales:
Spanish
Voice Aarav · AudioSonic Premium
Japanese
Voice Abby · AudioSonic Premium
How to write directions that actually land
A few rules get you dramatically better results:
- Put one direction at the very start of a segment. To change the delivery mid-script, start a new segment with a new direction.
- Write it in English, even for non-English text.
- Lowercase, no punctuation inside the bracket.
[speak softly and slowly]beats[Speak Softly, and Slowly.]. - Be descriptive. The more you describe the performance — mood, pace, pitch, vocal style — the more nuanced the result.
- Avoid contradictions. Do not ask for a whisper and a shout in the same breath.
Try it
Open the Text to Speech studio and direct your first line — it is free to start, no card required. When you are ready for unlimited custom voices and higher limits, the pricing starts at $20/mo for Creator.
The words are yours. Now the performance is too.
![Directing markup — an [excited] tag highlighted over a waveform on an indigo gradient](/cdn-cgi/image/width=2048,quality=75,format=auto,fit=scale-down/audiopod_og.png)