How to Transcribe Long Videos and Multi-Hour Podcasts
Long-form content such as webinars, lectures, and podcasts can run for hours, which makes manual transcription impractical. AI transcription handles this well, but there are practical limits and a few habits that make long recordings easier to manage. This guide covers how to approach files that run well beyond an hour.
Know the per-file limit
On JotVox, each AI transcription handles up to 90 minutes of audio or video per file. That covers most single episodes and talks. If your recording is longer than 90 minutes, you will need to split it into shorter segments and transcribe each one separately. This is a straightforward step, and it also keeps individual transcripts easier to navigate.
Splitting a longer recording
To transcribe a recording that exceeds the limit, divide it into parts before uploading. A few tips:
- Split at natural breaks, such as the end of a segment or a pause between speakers, so no sentence is cut in half.
- Keep each part under 90 minutes with a little margin to spare.
- Name the files in order so you can reassemble the transcripts cleanly afterward.
You can split audio in most free editors by exporting sections as separate files, then upload each part. JotVox transcribes uploaded files from your device, Google Drive, or Dropbox in addition to public links, so a locally split file works fine.
Use timestamps to navigate
Every JotVox transcript comes with timestamped lines, which is especially useful for long content. Timestamps let you jump to the moment a topic was discussed and stitch multi-part transcripts together in the right order. If you plan to publish or reference exact moments, see how to work with a YouTube transcript with timestamps.
Choose the right export
For a readable document, export as TXT. For captions or subtitle files aligned to the video, export as SRT or VTT. If you are unsure which fits your project, the guide on SRT vs VTT vs TXT compares them. You can also download the audio as MP3 and generate an optional AI summary in the source language, which is handy for quickly reviewing a long episode.
Plan for accuracy on long files
Long recordings often contain multiple speakers, changing audio conditions, and stretches of background noise. Accuracy can vary across a single file for those reasons, so budget time to review the draft. If the source already has captions, reusing them is free and can save credits on longer content. Note that JotVox does not store your files after processing, so keep your own copies of both the media and the finished transcripts.