Podcasts, interviews, lectures, calls and voice memos all end up as the same thing here: a timestamped transcript you can read, search and export. Upload the file or paste a public link.
Open JotVox →Upload the recording from your device, or select it in Google Drive or Dropbox so a long file does not have to be downloaded first. AI transcription reads the audio and returns a timestamped transcript at 98%+ accuracy across 99 languages, using credits.
If the audio is published with a public link, paste that instead. Where the platform already has captions, JotVox grabs them, which is free and needs no sign-up. Audio platforms include SoundCloud and Apple Podcasts.
Audio is processed in memory and deleted right after transcription.
The format matters much less than people expect, because speech survives ordinary compression well.
Per-format detail lives on MP3 to text, M4A to text and WAV to text.
There is no benefit to converting between these before uploading. Re-encoding cannot recover anything the original recording missed, and it costs you a step. Upload whatever came off the recorder.
Recording conditions, almost entirely. The same system produces a near-perfect transcript of a close-miked interview and a patchy one of a room recording, from identical settings.
More on this in what affects transcription accuracy.
Processing time tracks the runtime, so a long episode or a full day of interviews takes longer than a short clip, and a large upload adds transfer time before transcription starts.
Credits pay for AI transcription and AI summaries. Grabbing captions that already exist on a link costs nothing. Signup gives you 3 free credits and one free AI transcription per day after email verification, with credit packs from $4.99 for 20 credits and subscriptions from $9 per month.
The daily free transcription is enough for occasional use, such as one interview or one episode at a time. Regular batches of recordings are what the packs and subscriptions are for.
Every transcript is timestamped, which is what makes a long recording usable: you jump to the moment rather than scrolling for it.
Export TXT for reading and quoting, SRT or VTT to keep the timings as caption cues, and MP3 for a compact copy of the audio. AI summaries condense a recording into key takeaways, chapters, quotes or a timeline, and use credits.
Which export you want depends on where the text is going. Notes, articles and search indexes want TXT. Anything that will sit under a video wants SRT or VTT, because those keep the cue timings the player reads.
1. Add the recording Upload the audio from your device, or pick it out of Google Drive or Dropbox. If the same recording is published publicly, pasting its link works too and may be free.
2. Transcribe AI transcription reads the audio at 98%+ accuracy across 99 languages and uses credits. Where a link already carries captions, JotVox grabs those instead, with no sign-up needed.
3. Use the timestamps to review Jump to the passages that matter and check names, acronyms and any moment where speakers overlapped or the microphone was far from the talker.
4. Export or summarise Download TXT, SRT, VTT or MP3, or generate an AI summary with key takeaways, chapters, quotes or a timeline. Summaries use credits.
AI transcription runs at 98%+ accuracy across 99 languages on reasonably clear speech. The variable is the recording, not the language. Close microphones, quiet rooms and speakers who take turns produce transcripts that need almost no correction, while distant recording, echo, background noise and crosstalk are what push the error rate up on any system.
Common recording and distribution formats including MP3, M4A and WAV. MP3 is typical for podcasts, M4A for phone voice memos and meeting recorders, and WAV for editors and field recorders. There is no need to convert between them before uploading, since re-encoding cannot add back detail the original recording did not capture.
Yes. AI transcription covers 99 languages, so interviews, lectures and podcasts recorded in most widely spoken languages can be transcribed directly. This is also a good reason to run AI transcription rather than reuse an existing caption track, which will only ever exist in whatever languages the platform generated.
No. Audio is processed in memory and deleted right after transcription completes. Your transcript remains available to read and export as TXT, SRT, VTT or MP3, but the source audio itself is not retained afterwards. That applies to files uploaded from your device as well as those selected from Google Drive or Dropbox.