Two ways in: paste a link to a video that is already online, or upload the file from your device, Google Drive or Dropbox. Both end with a timestamped transcript you can export as TXT, SRT, VTT or MP3.
Open JotVox →They solve different problems.
Links are supported across 19+ platforms including TikTok, YouTube, Instagram, X, LinkedIn, Twitch and Bilibili. The full list is on platforms.
Only the audio track is transcribed. The picture is never the input, so resolution, frame rate and codec make no difference to the text you get back.
Upload from your device, or select the file in Google Drive or Dropbox so a large recording does not have to be downloaded first. If a particular file is not accepted, exporting or remuxing it to MP4 is the reliable fallback, since it is the most widely handled video container.
Audio is processed in memory and deleted right after transcription.
Existing captions are the fastest and cheapest path, but they are only as good as whoever or whatever made them. Platform auto-captions routinely miss names, product terms, acronyms and numbers, and they carry no speaker structure.
AI transcription runs on the audio at 98%+ accuracy across 99 languages and uses credits. It is the better option when the captions are auto-generated and rough, when there are none at all, or when you need the words in another language. That last case is covered in transcribing a video in any language.
A useful rule of thumb: if you are going to quote the video, publish the text, or hand it to someone else, transcribe the audio. If you just need to find where a topic came up, existing captions are enough.
Processing time scales roughly with runtime, and an upload adds transfer time before any of that starts. Webinars, lectures and full recordings simply take longer than clips.
Accuracy depends on the recording. Screen captures and single-speaker presentations transcribe very cleanly. Panels, handheld phone video shot across a room, and anything with heavy background noise or crosstalk will need a review pass.
Signup includes 3 free credits and one free AI transcription per day after email verification. Credit packs start at $4.99 for 20 credits, and subscriptions start at $9 per month.
TXT for reading, quoting and search. SRT and VTT when the text is going back onto the video as subtitles. MP3 when you only want the audio.
AI summaries condense a long transcript into key takeaways, chapters, quotes or a timeline, which uses credits. Chapters in particular are worth generating on anything over half an hour, because they give a long recording a table of contents it never had.
For a specific container start at MP4 to text. If the transcript is heading for a subtitle workflow, SRT to VTT explains how those two formats differ, and audio to text covers sound-only sources.
1. Paste a link or upload the video Use the link if the video is public on a supported platform. Otherwise upload the file from your device, Google Drive or Dropbox. Only the audio track is used either way.
2. Let JotVox pick the method Where a link already has captions, JotVox grabs them free with no sign-up. Where there are none, or they are unusable, AI transcription reads the audio at 98%+ accuracy in 99 languages and uses credits.
3. Work through the timestamped transcript Jump to any moment using the timestamps, check names, acronyms and numbers, and confirm passages where speakers overlapped or the microphone was distant.
4. Export or summarise Download TXT, SRT, VTT or MP3. Optionally generate an AI summary with key takeaways, chapters, quotes or a timeline, which uses credits.
Paste the video's link or upload the file to JotVox. If the link is on one of the 19+ supported platforms and already has captions, JotVox grabs them free with no sign-up. Otherwise AI transcription reads the audio track at 98%+ accuracy across 99 languages, using credits, and returns a timestamped transcript you can export as TXT, SRT, VTT or MP3.
Yes, that is what AI transcription is for. When a video has no caption track, or the platform never generated one, JotVox transcribes the audio itself rather than relying on anything the source provides. This works for uploaded files as well as links, covers 99 languages, and uses credits from your account.
Upload video files from your device, Google Drive or Dropbox and JotVox transcribes the audio track inside them. MP4 is the most universally handled container, so if a particular file is not accepted, exporting or remuxing it to MP4 first is the reliable fix. Video resolution and codec never affect transcript quality, only the audio does.
Yes, when the video is on a supported platform and already has captions. Grabbing those is free and requires no sign-up. For everything else, signing up gives you 3 free credits and, after email verification, one free AI transcription each day. Paid credit packs start at $4.99 for 20 credits and subscriptions start at $9 per month.