AudioPod AI
  • Pricing

Transcription

Editing and exporting a transcript

Fix the transcript before you export it, then pick the right file: subtitles for video, a document for people, JSON for anything a machine has to read.

Lesson 3 · core · 6 min read

Open Transcription

A transcript is finished when the words are right and it is in the file the next step needs. Those are two separate jobs and they belong in that order — exporting first and correcting afterwards means doing the export twice.

Fix it before you export it

Every segment opens for editing, and each one carries the same 4 fields. Edits are saved against the transcript itself, so a name you correct once is correct in every format you download afterwards.

Field
What it is for
Speaker label
Who this segment belongs to. Retype it to reassign a crossed segment, or to replace the generated label with a person's actual name.
Start
When this segment begins, in seconds. Worth nudging when a caption comes in slightly ahead of the line being spoken.
End
When it ends. The other half of a caption that sits a beat off the audio.
Text
The words themselves. Names, jargon and numbers are where a transcript is wrong, and they are usually wrong the same way every time.

What an open segment lets you change.

  1. 01

    Start with the repeated errors

    Transcripts are wrong systematically, not randomly. A surname, a product name, a piece of jargon — each is usually wrong the same way every time it appears, so a handful of fixes clears most of the errors in the file.

  2. 02

    Then the speakers

    Reassign the segments that crossed, and put real names where the generated labels are. This is per segment, so it is worth doing on a transcript you are going to publish and not worth doing on one you are going to skim.

  3. 03

    Timings last, and only if they matter

    Nudge a start or an end when a caption sits off the speech. In a document the timings are invisible, so this step is skippable for everything except subtitles.

Careful

A file you have already downloaded does not change when you edit the transcript. Correct first, export second — and if you have already sent the wrong version out, re-export rather than patching the file by hand.

Which format goes where

The download menu offers 7 formats, and the choice is entirely about the destination rather than about quality — each one is generated from the same saved transcript.

Format
Where it belongs
Text (.txt)
Just the words. The right choice when the transcript is going into something else — a draft, an email, a prompt — and nothing about timing or layout matters.
JSON (.json)
The structured version: segments, times, speakers, and per-word timings when word timestamps were on. This is the one to take if code is going to read it.
SubRip (.srt)
Subtitles. The format almost every video editor and video platform accepts, and the default answer for captions.
WebVTT (.vtt)
Subtitles for the web — what an HTML video player expects. Take this one when the captions are going onto a page you control rather than into an editor.
Word (.docx)
A document someone is going to read and mark up. The right export for an interview a human has to work through, not a machine.
PDF (.pdf)
A document nobody is going to edit — a record to file, send or attach. Fixed layout, so fix the transcript before you export it.
HTML (.html)
A web page you can open in a browser or paste into a CMS with its formatting intact.

Every export format, and where it belongs.

  • Going into a video: SubRip (.srt) and WebVTT (.vtt). SubRip is the safe default; WebVTT is for a web player.
  • Going to a person: Word (.docx), PDF (.pdf) and HTML (.html). Word if they will edit it, PDF if they will not.
  • Going into software: JSON (.json) — the only export that carries the structure, and the only one worth parsing.
  • Going into a draft: Text (.txt). Nothing but the words, which is the point.

Tip

Exporting is not a commitment. Download SubRip for the edit and Word for the write-up from the same transcript — the file is generated fresh each time from what is saved.

What this does not do

Worth knowing before you plan a workflow around it. AudioTranscribe turns speech into text and labels who was speaking; it does not translate, it does not burn subtitles into a video, and it does not rename a speaker across a whole transcript in one action.

  • Renaming a speaker is per segment. Declaring the speaker count before the job is the cheaper path — see the previous lesson.
  • A caption file is a file, not a video. Your editor or player applies it; nothing here writes it onto the picture.
  • The Link tab is YouTube-only, so anything hosted elsewhere has to be uploaded.

That is the track. If a transcript came out worse than you expected, the fix is almost never in this lesson — go back to accuracy and the speaker count, because those two decide what you are editing in the first place.

Questions

What people ask about this

Free to start

Now go make one

Reading about a style description only gets you so far. The studio is free to use — write one sentence and hear what comes back.

Create a free account

1,000 credits every month. No card required.

Previous lessonAccuracy, speakers and timestamps

All transcription lessons

Make something worth hearing.

Start creating free

Create

  • Music
  • Text to speech
  • Audiobooks
  • Podcasts
  • Voice changer
  • Audio reader
  • Narration

Edit & convert

  • Stem splitter
  • Separate speakers
  • Noise reduction
  • Speech to text
  • Media converter
  • Browser DAW
  • All features

Developers

  • Developer hub
  • API reference
  • Quickstart
  • Python SDK
  • MCP server
  • Changelog
  • API status

Resources

  • Guides
  • Languages
  • Use cases
  • Alternatives
  • Tool comparisons
  • AI audio guide
  • Glossary
  • Showcase

Free tools

  • Audio Format Converter
  • Video to Audio Extractor
  • Voice Recorder
  • Free Stem Splitter
  • Free Vocal Remover
  • All free tools

Company

  • About
  • Manifesto
  • Careers
  • Blog
  • Customers
  • Affiliate program
  • Contact
All pages · Sitemap

Studio

  • AI Music & Rap
  • Text to Speech
  • Audiobook Studio
  • Podcast Generator
  • Voice Changer
  • Audio Reader
  • AI Narrator
  • Studio overview
  • All features

Edit & process

  • Stem Splitter
  • Speaker Separation
  • Noise Reduction
  • Speech to Text
  • Media Converter
  • Browser DAW
  • YouTube to Podcast

Voices

  • Voice library
  • Languages
  • Iconic voices
  • Showcase
  • Music Radio

Free tools

  • All free tools
  • Audio Format Converter
  • Video to Audio
  • Audio Trimmer
  • Voice Recorder
  • ACX Checker
  • Free Stem Splitter
  • WAV to MP3 Converter
  • MP4 to MP3 Converter

Solutions

  • Audiobook authors
  • Podcasters
  • Musicians & creators
  • Education
  • Voice agents
  • Gaming
  • Accessibility
  • Advertising
  • All use cases
  • Authors program
  • Enterprise

Compare

  • vs ElevenLabs
  • vs Suno
  • vs Descript
  • vs Murf
  • vs NotebookLM
  • vs LALAL.AI
  • vs NarrationBox
  • All alternatives
  • Tool comparisons

Resources

  • Blog
  • Guides
  • Music Studio guides
  • Audiobook guides
  • Speaker Separation guides
  • Stem Splitter guides
  • Voice Studio guides
  • Transcription guides
  • Changelog
  • Launches
  • Customers
  • Glossary
  • AI Audio guide
  • Family voice (mobile)
  • AudioPod mobile
  • Affiliate program
  • Pricing
  • Developers
  • For AI agents
  • AudioPod for Startups

Company & legal

  • About
  • Manifesto
  • Careers
  • Press & media
  • Contact
  • Responsible AI
  • Voice consent
  • Trust & security
  • System status
  • Security disclosures
  • Security policy
  • Privacy
  • Cookie policy
  • Terms
AudioPod AI

© 2026 AudioPod AI, Inc. All rights reserved.

Privacy|Terms|Trust Center|Responsible AI|Voice consent