A WebVTT file carries a header line, optional cue identifiers, cue settings, comment and styling blocks, and inline speaker tags. All of that has to come out before you have readable text. Here is what to remove, and when transcribing the source again is the better move.
Open JotVox →WebVTT (.vtt) is plain text like SRT, but it holds more than caption lines. A file can look like this.
WEBVTT NOTE exported from the edit intro 00:00:02.140 --> 00:00:05.600 align:start position:10% <v Priya>So the first thing we did was rewrite the ingest pipeline.
The first line is always WEBVTT. Cues may have a text identifier above the timestamp instead of a number. The time range uses a period before the milliseconds, and the hours field can be left off entirely.
Anything after the arrow and end time is a cue setting for placement or alignment. NOTE and STYLE blocks and inline tags such as the voice span are also part of the format.
Work from the top of the file down.
A text editor with regular expression search does most of this, and a short script does all of it. Nothing here is a conversion; you are deleting the parts of a text file that describe timing and display.
Cue text is broken to fit a video frame for a second or two. Once the timestamps are gone, those breaks remain and the result reads as fragments rather than sentences. Speaker changes vanish too, unless the file used voice tags and you turned them into name prefixes on the way through.
If the captions were auto-generated, the errors survive the strip untouched. Reformatting never improves the words themselves.
JotVox is a transcription tool, not a subtitle file converter. It will not take a VTT off your disk and reformat it. Use the steps above for that.
Give JotVox the source instead, either as a pasted link or an uploaded video or audio file, and it produces a timestamped transcript you can export as TXT, SRT, VTT or MP3. That is worth doing when the VTT is missing, machine-generated and rough, in the wrong language, or when you want prose with paragraphs rather than caption cues.
Files on disk are covered on video to text, hosted video on Vidyard transcripts.
The TXT export is continuous readable text, suited to notes, search and quoting. SRT and VTT exports keep the timings if you still need cues for a player.
Existing captions grabbed from a link are free with no sign-up. AI transcription, for audio with no usable captions, runs at 98%+ accuracy in 99 languages and uses credits. You get 3 free credits at signup and one free AI transcription a day after email verification. Audio is processed in memory and deleted right after transcription.
For the SRT equivalent of this page, see SRT to TXT, or VTT to SRT to move between the two caption formats.
1. Remove the header and metadata blocks Delete the opening WEBVTT line, then any NOTE, STYLE or REGION blocks. None of them contain caption text, and they will otherwise land in your transcript.
2. Delete timestamp and identifier lines Remove every line containing the time range arrow, including cue settings such as align or position that follow the end time, plus any cue identifier line sitting directly above it.
3. Clean the remaining cue text Strip inline tags such as voice, bold and italic spans while keeping the words inside, convert character escapes back to plain characters, then rejoin the fragments into sentences.
4. Or transcribe the source If the VTT is missing or the captions in it are poor, paste the link or upload the media to JotVox and export the transcript straight to TXT.
A VTT is a caption file, so it describes when and where text appears on a video. It begins with a WEBVTT header, splits speech into short timed cues, and can carry placement settings, styling blocks and speaker tags. A transcript is just the words, written as sentences and paragraphs for reading rather than for display over a picture.
No. JotVox transcribes video and audio; it does not accept a subtitle file as input and reformat it. It exports finished transcripts as TXT, SRT, VTT or MP3. When your existing VTT is rough or in the wrong language and you still have the source video, transcribing it again usually beats trying to salvage the captions.
Yes, if you want a clean transcript. The header is required at the top of a valid WebVTT file, but it is not caption text, so it would sit as a stray word at the start of your TXT. The same applies to NOTE comments, STYLE blocks and REGION definitions, which describe presentation rather than speech.
WebVTT allows the hours field to be omitted, so a cue can read 02:05.600 rather than 00:02:05.600. It is valid either way. This matters if you plan to move the file to SRT later, since SRT conventionally uses the full hours, minutes, seconds and milliseconds form and some players are strict about it.