Upload the MP4 or paste the link and JotVox returns a timestamped transcript you can read, search and export as TXT, SRT, VTT or MP3. Recordings, interviews, lectures, webinars and screen captures all work the same way.
Open JotVox →An MP4 is a container holding a video track and an audio track. Only the audio matters for a transcript, so the picture, resolution and codec make no difference to the result.
Two routes into JotVox:
Either way you get a timestamped transcript in the app, ready to export.
If the video is public on a supported platform, the link is faster and often free, because existing captions can be pulled without transcribing anything. That covers 19+ platforms including YouTube, TikTok, Instagram, X and LinkedIn.
Upload the file when the MP4 is private, local, or has no captions worth using. Auto-captions frequently mangle names, product terms and numbers, and an AI transcription of the audio is usually the cleaner starting point.
Files in cloud storage do not need downloading first; connect Google Drive or Dropbox and pick the MP4 there.
AI transcription runs at 98%+ accuracy across 99 languages. That figure assumes reasonably clear audio, and the things that pull it down are predictable.
Screen recordings and single-speaker webinars usually transcribe very cleanly. A handheld recording of a busy panel will need more review.
None of this is fixed by the file. Re-encoding an MP4 to a higher bitrate cannot add back speech the microphone never picked up clearly, so the useful effort goes into the recording, not the export settings.
Long videos take longer to process than short ones, roughly in proportion to their runtime, and a large upload also has to travel to the server first. A two hour recording is a longer wait than a five minute clip on both counts.
AI transcription uses credits. Signup includes 3 free credits and one free AI transcription per day after email verification. Credit packs start at $4.99 for 20 credits and subscriptions start at $9 per month. Pulling existing captions from a link stays free.
Audio is processed in memory and deleted right after transcription.
TXT for reading and quoting, SRT and VTT when the text is going back onto video as subtitles, and MP3 when you want the audio on its own. AI summaries can turn a long transcript into key takeaways, chapters, quotes or a timeline, which uses credits.
The transcript stays timestamped in the app, so a two hour webinar is navigable rather than a wall of text. That is usually more valuable than the export itself, because it lets you find the ninety seconds you actually needed.
If you need the subtitle formats specifically, SRT to VTT covers how they differ. For audio-only sources, see MP3 to text, and video to text for other video formats.
1. Add the video Upload the MP4 from your device, Google Drive or Dropbox. If the same video is already online on a supported platform, paste its link instead and skip the upload.
2. Choose captions or AI transcription When a pasted link already has captions, JotVox grabs them free with no sign-up. When there are none, or they are poor, AI transcription reads the audio track and uses credits.
3. Review the timestamped transcript The transcript appears with timestamps so you can jump to any moment, check names and technical terms, and confirm the wording before exporting.
4. Export the format you need Download as TXT for plain text, SRT or VTT for subtitles, or MP3 for the audio. Optional AI summaries produce takeaways, chapters, quotes or a timeline.
Upload the MP4 to JotVox from your device, Google Drive or Dropbox, or paste the link if the video is already online. JotVox reads the audio track and returns a timestamped transcript. Existing captions from a supported link are free with no sign-up, while AI transcription of the audio uses credits and runs at 98%+ accuracy across 99 languages.
No, but audio quality does. An MP4 holds separate video and audio tracks, and only the audio is transcribed, so resolution, frame rate and codec are irrelevant. What matters is how clearly the speech was captured. Close microphones and low background noise produce clean transcripts; distant recording, echo and overlapping speakers cause most errors.
Yes. Uploading is the route for anything not published on a platform. Send the file from your device, or pick it directly from Google Drive or Dropbox without downloading it first. AI transcription then runs on the audio and uses credits. The audio is processed in memory and deleted right after transcription.
Processing time scales roughly with the length of the recording, so a short clip finishes quickly and a multi-hour session takes noticeably longer. Large uploads also spend time transferring before any transcription starts. Pasting a link where captions already exist is the fastest option, since nothing has to be uploaded or transcribed.