Upload an MP3 and get a timestamped transcript back, exportable as TXT, SRT or VTT. Podcasts, interviews, calls, lectures and voice notes are the usual sources, and each has its own quirks worth knowing about.
Open JotVox →Upload the file from your device, or pick it straight out of Google Drive or Dropbox without downloading it first. JotVox transcribes the audio with AI and returns a timestamped transcript, which uses credits.
If the same recording is published somewhere with a public link, paste that instead. When captions already exist on the platform, JotVox grabs them for free with no sign-up. Podcast and audio platforms are covered on Apple Podcasts transcripts and across all supported platforms.
AI transcription runs at 98%+ accuracy across 99 languages on reasonably clear speech. MP3 is a lossy format, but at ordinary podcast and voice-recorder bitrates that compression is not what limits the result. Recording conditions are.
Very low bitrate mono recordings, the kind some call recorders produce, are the one place where the encoding itself starts to hurt.
Re-encoding an MP3 to a higher bitrate does not help. The detail was discarded when the file was first made, and nothing downstream can put it back. If you still have a better master of the same recording, transcribe that instead.
Processing time tracks the length of the recording, so a ninety minute episode takes meaningfully longer than a ten minute segment, and a large file also has to upload before anything begins.
For long audio, the timestamps do most of the work afterwards. They let you jump to a moment rather than scrolling, which is the difference between a transcript you use and one you archive. AI summaries can also compress an episode into key takeaways, chapters, quotes or a timeline, and that uses credits.
More on long-form recordings in transcribing long videos and podcasts.
Credits pay for AI transcription and AI summaries. Grabbing captions that already exist on a link costs nothing and needs no account.
Signup gives you 3 free credits, plus one free AI transcription each day once your email is verified. Credit packs start at $4.99 for 20 credits, and subscriptions start at $9 per month for regular use.
Audio is processed in memory and deleted right after transcription.
Export as TXT for reading, quoting and search. SRT and VTT are there when the audio is paired with video and you need timed captions rather than prose. MP3 export is useful when you started from a video and want just the sound.
Show notes, quote pulls and searchable archives all come out of the TXT export. If the transcript is going to feed another tool, TXT is the format to reach for, since it carries no cue numbering or timing syntax to strip out first.
Other audio formats work the same way; see M4A to text, WAV to text or the broader audio to text page.
1. Upload the MP3 Send the file from your device, or select it directly in Google Drive or Dropbox so you do not have to download it first. A public link to the same recording also works.
2. Run the transcription AI transcription reads the audio and uses credits. If you pasted a link to a platform that already has captions, JotVox grabs those instead, free and without a sign-up.
3. Check names and terms Use the timestamps to jump to the parts that matter and confirm proper nouns, acronyms and numbers, which are the words most often misheard in any recording.
4. Export or summarise Download TXT, SRT or VTT, or generate an AI summary with key takeaways, chapters, quotes or a timeline. Summaries use credits.
JotVox transcribes at 98%+ accuracy across 99 languages on reasonably clear speech. MP3 compression itself rarely limits the outcome at normal podcast or recorder bitrates. Recording conditions do: separate microphones, low background noise and speakers who do not talk over each other produce the cleanest results, while echo and crosstalk cause most of the errors you will see.
Yes. You can pick the file directly from Google Drive or Dropbox rather than downloading it to your device and uploading it again. That saves the round trip on large episodes. AI transcription then runs on the audio and uses credits, and the audio is processed in memory and deleted right after transcription.
Credits cover AI transcription and AI summaries. Pulling captions that already exist on a supported link is free and needs no sign-up. New accounts get 3 free credits and, after email verification, one free AI transcription per day. Credit packs start at $4.99 for 20 credits and subscriptions start at $9 per month.
Yes. Transcripts are timestamped, so you can jump straight to a moment in a long episode instead of scrolling through text. If you need those timings in a file, export SRT or VTT, which keep them as cues. The TXT export is the continuous reading version without cue formatting.