JJotVox Open the app →

Caption vs Transcript vs Subtitle: What's the Difference?

MCBy Maya Chen · Jul 28, 2026 · 6 min read

The words caption, transcript, and subtitle are often used interchangeably, but they describe three distinct things. Each has a different format, a different purpose, and a different intended audience. Knowing how they differ makes it easier to choose the right output when you work with video and audio.

What a transcript is

A transcript is the full text of everything spoken in a recording, written out as a continuous document. It is not tied to the video timeline in the way captions are, though it may include timestamps that mark when each line was said. Transcripts are commonly used for reading, searching, quoting, and repurposing content. Because they are plain text, they are easy to edit, copy into other documents, or index for search.

What captions are

Captions are short blocks of text displayed on screen, synchronized to the audio, and read by viewers as the video plays. They were designed primarily for accessibility, so that people who are deaf or hard of hearing can follow along. Captions are usually in the same language as the audio and often include non-speech cues such as [music playing] or [applause]. Technically, captions are stored as timed text files that pair each line with a start and end time.

What subtitles are

Subtitles are also timed on-screen text, but their traditional purpose is translation. They assume the viewer can hear the audio but may not understand the language, so they typically render dialogue in a different language and leave out non-speech sounds. In everyday usage the line between subtitles and captions has blurred, and many platforms use the terms loosely.

A quick comparison

File formats overlap. Captions and subtitles are commonly stored as SRT or VTT files, while a plain transcript is often a TXT file. The same underlying words can be exported in any of these forms depending on what you need.

Which one do you need?

If you want something to read, search, or reuse as written content, a transcript is the natural choice. If you want text that appears over the video for accessibility, you want captions. If your audience speaks a different language, subtitles are the relevant format. With JotVox you can produce a timestamped transcript and export it as TXT, SRT, or VTT, so you can move between these uses without redoing the work. If your source has no on-screen text at all, you can still get text from a video with no captions using AI transcription.

Try it: paste a video link into JotVox — grabbing existing captions is free, and new accounts get 3 free credits plus one free AI transcription a day.

None of these formats is inherently better than the others. They are tools for different jobs, and most video projects end up using more than one.

Frequently asked questions

Are captions and subtitles the same thing?+
Not originally. Captions were designed for accessibility and are usually in the same language as the audio, often noting non-speech sounds. Subtitles were designed for translation and typically leave sounds out. In casual use the terms are now often mixed.
Is a transcript just captions without timing?+
Roughly, yes. A transcript is the full spoken text as a readable document. Captions take similar text and split it into short timed blocks synchronized to the video. A transcript may still include timestamps, but it is meant to be read on its own.
Can I turn a transcript into captions?+
Yes. If the transcript includes timing information, it can be exported as a caption file such as SRT or VTT. JotVox lets you export the same transcript as TXT, SRT, or VTT so you can use it as a document or as captions.
Try JotVox free →

Keep reading

How to Transcribe a TikTok Video (2026) — Free, With Timestamps
How to Get an Instagram Reel Transcript (2026 Guide)
How to Get a YouTube Transcript With Timestamps (2026)