JJotVox Open the app →
Subtitle formats

SRT to VTT

Moving an SRT to WebVTT is a small, well defined edit: add a header, swap the millisecond separator, and decide what to do with the cue numbers. The parts that catch people out are encoding and how browsers read cue text.

Open JotVox →

What actually changes between SRT and VTT?

The two formats hold the same thing, timed lines of text, but differ in a handful of concrete ways.

SubRip
1
00:00:02,140 --> 00:00:05,600
Rewriting the ingest pipeline.

WebVTT
WEBVTT

1
00:00:02.140 --> 00:00:05.600
Rewriting the ingest pipeline.

A WebVTT file must open with the line WEBVTT, followed by a blank line. Timestamps use a period before the milliseconds where SRT uses a comma. SRT numbers its cues; in WebVTT that line is an optional identifier, so it can stay or go.

WebVTT also adds things SRT has no equivalent for: cue settings for placement after the end time, NOTE comments, STYLE blocks and voice tags. Converting up from SRT simply leaves those unused.

How to convert an SRT to WebVTT by hand

That is the whole edit. It is deliberately small, which is why plain text tools handle it well.

What breaks after the conversion?

Three things account for most failures when a converted file will not play in a browser.

Encoding. WebVTT expects UTF-8. An SRT saved in an older regional encoding will show mangled accented characters, and a stray byte order mark at the top can stop the header being recognised.

Markup characters. WebVTT treats the less-than sign as the start of a tag and the ampersand as the start of an escape. Caption text containing either should be escaped, or the player may swallow part of the line.

Serving. A .vtt loaded by a video element has to be served as text/vtt and, if it sits on another origin, with the cross-origin headers that permit it.

Why do you need VTT rather than SRT?

Usually because the captions are going onto a web page. WebVTT is the format the HTML video element reads through a track element, so browser playback is its home ground. SRT stayed dominant on desktop players and in a lot of upload pipelines, which is why files arrive in it.

If you are choosing between the two rather than converting, SRT vs VTT vs TXT lays out the tradeoffs.

Skipping the conversion with JotVox

JotVox is a transcription tool, so it does not take a subtitle file in and hand a different one back. It goes the other way: paste a link or upload the media, get a timestamped transcript, then export it as VTT directly. No conversion step exists to get wrong.

That helps when the SRT you were about to convert is a poor auto-caption, is in the wrong language, or does not exist. Grabbing existing captions from a link is free with no sign-up. AI transcription for audio with no captions runs at 98%+ accuracy in 99 languages and uses credits.

Also see VTT to SRT for the reverse edit, MP4 to text for a video file, or YouTube transcripts for a link.

How to convert an SRT file to WebVTT

1. Add the WEBVTT header Open the .srt in a text editor and insert WEBVTT as the very first line, followed by a blank line before the first cue. A WebVTT file without that header will not be recognised.

2. Change the millisecond separator On every timestamp line, replace the comma before the milliseconds with a period, so 00:00:02,140 becomes 00:00:02.140. Restrict the replacement to timestamp lines so caption punctuation is untouched.

3. Save as UTF-8 with a .vtt extension WebVTT expects UTF-8. Re-save in that encoding, avoid a byte order mark before the header, and rename the file with the .vtt extension.

4. Or export VTT from the transcript If you still have the video or its link, run it through JotVox and export the transcript as VTT directly, which skips the conversion entirely.

Frequently asked questions

What is the difference between SRT and VTT?

WebVTT starts with a required WEBVTT header line and uses a period before the milliseconds, while SRT uses a comma and numbers each cue. WebVTT also supports cue settings for on-screen placement, comment and styling blocks, and speaker voice tags, none of which SRT has. SRT is common in desktop players and upload pipelines; WebVTT is what browsers read.

Can I just rename an .srt file to .vtt?

No. The extension is not what makes the file valid. Without the WEBVTT header line at the top, a browser will reject it, and timestamps still using commas will not parse as cue times. The rename works only after you have added the header and switched the millisecond separator to a period.

Do I have to keep the cue numbers?

No, they are optional in WebVTT. SRT requires a sequential index above each timestamp, but WebVTT treats that line as an optional cue identifier, and it does not have to be a number. Leaving the numbers in place is harmless and makes the file easier to diff against the original, so most conversions simply keep them.

Can JotVox give me a VTT file without converting anything?

Yes. JotVox produces the transcript from the source media rather than from a subtitle file, and VTT is one of its export formats alongside TXT, SRT and MP3. Paste a video or audio link, or upload a file from your device, Google Drive or Dropbox, then export the timestamped result as VTT.

Related