Caption vs Transcript vs Subtitle: What's the Difference?
The words caption, transcript, and subtitle are often used interchangeably, but they describe three distinct things. Each has a different format, a different purpose, and a different intended audience. Knowing how they differ makes it easier to choose the right output when you work with video and audio.
What a transcript is
A transcript is the full text of everything spoken in a recording, written out as a continuous document. It is not tied to the video timeline in the way captions are, though it may include timestamps that mark when each line was said. Transcripts are commonly used for reading, searching, quoting, and repurposing content. Because they are plain text, they are easy to edit, copy into other documents, or index for search.
What captions are
Captions are short blocks of text displayed on screen, synchronized to the audio, and read by viewers as the video plays. They were designed primarily for accessibility, so that people who are deaf or hard of hearing can follow along. Captions are usually in the same language as the audio and often include non-speech cues such as [music playing] or [applause]. Technically, captions are stored as timed text files that pair each line with a start and end time.
What subtitles are
Subtitles are also timed on-screen text, but their traditional purpose is translation. They assume the viewer can hear the audio but may not understand the language, so they typically render dialogue in a different language and leave out non-speech sounds. In everyday usage the line between subtitles and captions has blurred, and many platforms use the terms loosely.
A quick comparison
- Transcript: full text of the spoken content, read separately from the video, may include timestamps.
- Captions: timed on-screen text in the source language, built for accessibility, may note sounds.
- Subtitles: timed on-screen text, traditionally a translation for viewers who can hear but not understand the language.
File formats overlap. Captions and subtitles are commonly stored as SRT or VTT files, while a plain transcript is often a TXT file. The same underlying words can be exported in any of these forms depending on what you need.
Which one do you need?
If you want something to read, search, or reuse as written content, a transcript is the natural choice. If you want text that appears over the video for accessibility, you want captions. If your audience speaks a different language, subtitles are the relevant format. With JotVox you can produce a timestamped transcript and export it as TXT, SRT, or VTT, so you can move between these uses without redoing the work. If your source has no on-screen text at all, you can still get text from a video with no captions using AI transcription.
None of these formats is inherently better than the others. They are tools for different jobs, and most video projects end up using more than one.