AudioPod AI
  • Pricing

Transcription

Accuracy, speakers and timestamps

The three settings that decide whether a transcript is usable: when premium accuracy is worth paying for, how to get speaker labels you can trust, and what word timestamps unlock.

Lesson 2 · core · 7 min read

Open Transcription

Three settings do almost all the work here. One decides how well the words are heard, one decides whether the transcript knows who is talking, and one decides whether it knows exactly when. Everything else in the panel is a detail by comparison.

Accuracy

AudioTranscribe offers two accuracy tiers, and the honest way to choose is to judge the RECORDING rather than the importance of the job. A clear voice on a decent microphone comes back much the same either way; a noisy four-way call does not.

Tier
What it is for
Standard
Fast, and accurate for most audio: a clear single speaker, a decent microphone, ordinary vocabulary. Available on every plan.
Premium
Higher accuracy, better punctuation and proper nouns. Worth it for interviews, names, accents and tougher recordings. Available on Creator and above, from $20/mo.

The two tiers, and who can select them.

Write this
Not this
One person, quiet room, headset mic
Standard — the gap between the tiers has almost nothing to bite on here
Paying for Premium out of caution
Interview, two mics, some crosstalk
Premium — this is the case it exists for
Standard, then an hour repairing names
Conference call recorded through a laptop
Premium once, with the speaker count declared
Re-running the same tier hoping for a better result
A transcript that will be read, not published
Standard — good enough to search and skim is genuinely good enough
Premium on everything by default

One person, quiet room, headset mic

Write this:
Standard — the gap between the tiers has almost nothing to bite on here
Not this:
Paying for Premium out of caution

Interview, two mics, some crosstalk

Write this:
Premium — this is the case it exists for
Not this:
Standard, then an hour repairing names

Conference call recorded through a laptop

Write this:
Premium once, with the speaker count declared
Not this:
Re-running the same tier hoping for a better result

A transcript that will be read, not published

Write this:
Standard — good enough to search and skim is genuinely good enough
Not this:
Premium on everything by default

Note

Premium is a plan feature: Available on Creator and above, from $20/mo. On a free account the option is visible with a lock rather than hidden, so you can see what you are choosing between — and Standard is available to everyone, on every plan.

Speaker Diarization

Splits the transcript by who is talking and labels each segment with a speaker. On by default. Turn it off for a recording with one voice — there is nothing to separate, and a single speaker occasionally gets split in two. It is the setting people notice most, because a transcript that labels the wrong person is more annoying than a transcript that misspells a word.

Declares how many people are in the recording instead of letting it work that out. Only appears while diarization is on. Auto-detect, or an exact count from the picker — and if you know the number, say it. Separation is a judgement about how many distinct voices are present, and removing that judgement removes the two failure modes that come with it.

  • One person split into two labels: usually a change of tone, a moved microphone, or a stretch of worse-quality audio partway through.
  • Two people merged into one: usually similar voices, or one of them speaking only briefly.
  • Both are fixed by declaring the exact count — and neither is fixed by re-running with the same settings.

Tip

Turn separation OFF for a single-voice recording. There is nothing to split, and the only thing separation can do on a solo dictation is invent a second speaker.

Careful

Do not declare a count you are guessing at. Auto-detect being wrong is recoverable by editing a few segments; a confidently wrong count reshapes the whole transcript around a number that was never true.

Word Timestamps

Records a start and end time for every word rather than for each segment. This is what makes tight captions and word-level search possible later. They change nothing about the words you read on screen, which is why they are easy to dismiss — the difference appears in what you can do with the transcript afterwards.

  • Captions that break where the speech breaks rather than where a block of text happens to end — SubRip (.srt), WebVTT (.vtt) and JSON (.json) all carry timing.
  • Jumping to the exact moment a word was said, instead of to the top of the paragraph containing it.
  • Lining the transcript up against the audio in other software, which needs word-level times to do anything precise.

The next lesson is about what happens after all this: fixing what is left by hand, and choosing the export format that suits where the transcript is going.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonYour first transcriptNext lessonEditing and exporting a transcript

All transcription lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent