AudioPod AI
  • Pricing

Transcription

Your first transcript

Audio in, words out. Upload a file or paste a link, run it, and read the result back — one success before you touch a single setting.

Lesson 1 · start · 4 min read

Open Transcription

Transcription is the one tool here that usually works on the first attempt, which is why this lesson is short. Get one transcript out with everything left alone, then read the next lesson to find out which of the settings you skipped were actually costing you something.

Getting the audio in

There are 3 ways to hand over audio today, and a fourth tab that is not doing the work yet. Start with a file — it is the only route that does not depend on somebody else's site staying up.

Tab
What it takes
Upload
An audio or video file from your machine. The most reliable route, and the only one that does not depend on somebody else's site staying up.
Link
A YouTube URL. This tab takes YouTube links only — for audio hosted anywhere else, download it first and use Upload.
Live Mic
Transcribes as you speak, from your microphone. Good for a note or a short dictation; a recording you upload afterwards will read better than a live pass over the same words.
Live Captions
Captioning a meeting or another tab's audio. The tab exists but the feature does not yet — it currently points you at Live Mic.

The tabs at the top of the tool.

Careful

The Link tab is YouTube-only. A podcast RSS link, a shared drive URL or a direct file link will not work there — download the audio and upload it instead.

Run it with the defaults

The defaults are chosen to be right for the common case. Language is on Auto-detect, Accuracy is on Standard, and both Speaker Diarization and Word Timestamps start on. That combination gets you a labelled, word-timed transcript of most recordings without a single decision.

  • Standard accuracy: Fast, and accurate for most audio: a clear single speaker, a decent microphone, ordinary vocabulary.
  • Speaker Diarization on: the transcript comes back split by who is talking, rather than as one undivided block.
  • Word Timestamps on: every word carries its own start and end, which is what captions and search need later.

Note

The estimate you see before pressing the button is exactly that — an estimate, computed from the length of the audio. The charge is settled by the server when the job completes.

Read it back

Open the transcript and follow a minute of it against the audio. You are looking for three specific things, and each one points at a different setting in the next lesson.

  1. 01

    Are the words right?

    Names, jargon, numbers and accents are where a transcript goes wrong. If ordinary sentences are fine but every proper noun is mangled, that is the Accuracy control talking.

  2. 02

    Are the speakers right?

    Two people merged into one, or one person split across two labels, is the diarization guessing. You can tell it how many voices to expect — the next lesson covers when that helps.

  3. 03

    Do the times line up?

    If you are going to make captions, check a few segment boundaries against the audio now rather than after you have exported an SRT.

That is a transcript. It lives in your history, it is editable, and it exports in 7 formats — but none of that is worth doing until the words underneath are right, which is the next lesson.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Next lessonAccuracy, speakers and timestamps

All transcription lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent