WebVTT can express things SRT cannot, so this direction loses information rather than adding it. Cue settings, styling and speaker tags have to go somewhere, and every cue needs a number. Here is the complete edit list.
Open JotVox →Going up from SRT to WebVTT is mostly additive. Coming back down is lossy, because WebVTT carries features SRT has no slot for.
WEBVTT NOTE from the edit session intro 00:02.140 --> 00:05.600 align:start position:10% <v Priya>Rewriting the ingest pipeline.
Everything above except the words in the last line has to be resolved. The header goes, the NOTE block goes, the identifier intro becomes a number, the placement settings are dropped, the missing hours field is filled in, the period becomes a comma, and the voice tag either becomes a name prefix or disappears.
The failures are usually small and specific. A cue setting left after the end time makes the whole timestamp line unreadable to a strict parser, so the cue is skipped. Timestamps missing the hours field are the same story.
Numbering matters too. SRT expects an index above each timestamp, in order, and files that skip or repeat numbers can load partially in one player and not at all in another. Styling that survived from the WebVTT side is likely to appear as literal tag text on screen.
Save as UTF-8 unless a specific tool has told you otherwise, and check accented characters before you ship the file.
JotVox does not take a VTT in and return an SRT. It makes transcripts from media, then exports them as TXT, SRT, VTT or MP3.
So it is the better route when the conversion is not really your problem. If the WebVTT came from automatic captions and the names and technical terms are wrong, if you need the subtitles in another language, or if you only ever had the video and someone else's captions, starting from the audio produces a cleaner file than any amount of line editing.
See video to text for an uploaded file, or Dailymotion transcripts for a hosted link.
Paste a link or upload a video or audio file from your device, Google Drive or Dropbox. If the link already has captions, JotVox grabs them, which is free and needs no sign-up. If there are none, AI transcription runs on the audio at 98%+ accuracy across 99 languages and uses credits.
Export the result as SRT and the numbering, comma separators and cue spacing are already correct. Signup includes 3 free credits and one free AI transcription a day after email verification, with credit packs from $4.99 and subscriptions from $9 per month.
Related edits are covered on SRT to VTT and VTT to TXT.
1. Strip the WebVTT-only blocks Delete the WEBVTT header line and every NOTE, STYLE and REGION block. SRT has no equivalent for any of them, and left in place they appear as stray text in players.
2. Renumber the cues Replace any text cue identifiers with sequential numbers starting at 1, and add a number above each timestamp line that had none. SRT expects an index on every cue, in order.
3. Fix the timestamp lines Change the period before the milliseconds to a comma, pad timestamps that omit the hours field, and delete cue settings such as align or position that follow the end time.
4. Or export SRT from a transcript If the captions themselves are the weak link, transcribe the source with JotVox instead and export SRT directly, correctly numbered and formatted.
Positioning and styling. WebVTT cue settings such as align, position, line and size have no SRT equivalent, so on-screen placement reverts to the player default. STYLE blocks, REGION definitions and NOTE comments are dropped entirely, and speaker voice tags either become a plain name prefix in the text or disappear. The words and their timings survive intact.
Usually a timestamp line a player could not parse. Cue settings left after the end time, a period instead of a comma before the milliseconds, or a timestamp missing the hours field will each make a parser skip that cue while the rest of the file plays. Broken or repeated cue numbering causes the same kind of partial load.
Yes. WebVTT cue identifiers are optional and can be words, but SRT expects a sequential number above each timestamp line, starting at 1 and rising with no gaps. If your VTT used names like intro or no identifiers at all, add the numbering as part of the conversion or some players will reject the file.
No, not as a file conversion. JotVox transcribes video and audio and exports the result as TXT, SRT, VTT or MP3, so it never takes a subtitle file as its input. If you have the original media, transcribing it and exporting SRT gives you a correctly formatted file and, when the old captions were auto-generated, better text as well.