Long-form drama, variety and interview content turned into text you can search, quote and skim.
Get a Tencent Video transcript →Tencent Video (腾讯视频) is a licensed streaming service, not a creator feed. The catalogue is professionally produced drama, variety, film, documentary and interview programming, most of it running 40 minutes an episode and often much longer for variety formats. That changes the shape of the job entirely.
Professionally produced Chinese content usually ships with Chinese subtitles as part of the presentation, and a lot of that text is rendered into the video itself rather than carried as a separate readable track. Where a genuine caption track is available on a link, JotVox will take it, free and without an account. Where there is not one, AI transcription works from the audio instead.
A 90-minute variety episode is a different task from a three-minute clip. The raw transcript will run to tens of thousands of characters, which is more than anyone wants to read straight through.
That is what the AI summary layer is for. On top of the full text you can generate key takeaways, chapters that break the episode into segments, notable quotes, and a timeline. For media research the usual pattern is to summarize first, decide which ten minutes matter, then read that stretch of the full transcript closely.
You can also export the audio on its own as MP3 if you would rather listen to an interview at speed than read it. Working with long videos and podcasts goes deeper on the approach.
Drama dialogue is comparatively easy: scripted, recorded cleanly, one person speaking at a time. Variety is the opposite, and it is where speech recognition earns its keep:
JotVox's AI transcription is rated at 98%+ accuracy across 99 languages. A clean two-person interview lands near that. A shouting round of a game show does not, and you should expect to check names and puns by hand. Historical documentaries and studio interviews, the two formats media researchers use most, transcribe very well.
This is the honest constraint on Tencent Video more than on any other platform we support. Much of the catalogue is licensed content, which brings restrictions JotVox cannot work around:
Free previews, trailers, clips, promotional segments and interview extracts generally work fine. And if you have legitimate access to a file already, uploading it directly from your device, Google Drive or Dropbox skips the link problem entirely.
The audience here is research-shaped rather than casual:
The same workflow applies to other long-form catalogues: see Bilibili for lectures and creator long-form, Dailymotion for broadcast clips, and Niconico for the Japanese side.
The transcript is written in the language spoken on screen. A Mandarin drama comes back as Chinese text. JotVox supports 99 languages for AI transcription but does not translate between them, so there is no English-dub button here.
For subtitling and research that is the correct starting point: you want an accurate source-language transcript with timings, then translate from a solid base. Export SRT or VTT for the timed version, TXT for the reading version. Captions, transcripts and subtitles explains which of those three you are actually asking for.
1. Copy the episode page link Open the episode or clip in a browser and copy the address of that specific page. A link to the series landing page will not identify which episode you mean.
2. Check it plays without signing in If the page asks for a VIP subscription or a login before it will play, JotVox cannot reach the audio either. Use a free clip or preview, or upload a file you already have.
3. Paste it into JotVox and transcribe Drop the link into the box on the home page. An existing caption track is returned free; otherwise AI transcription runs on the episode audio.
4. Summarize before you read For a long episode, generate chapters and takeaways first to find the segments worth reading, then export the full text as TXT, SRT, VTT or MP3.
No. If an episode requires a paid subscription to play, the audio is not accessible to JotVox and there is nothing to transcribe. Free episodes, trailers, clips and promotional segments work normally. If you subscribe and can download or record content you are entitled to, you can upload that file to JotVox directly instead of using a link.
It transcribes the whole thing, then gives you tools to avoid reading all of it. AI summaries produce chapters, key takeaways, quotes and a timeline for the episode, so you can locate the ten minutes that matter and read only that part of the full transcript. Long episodes are also exportable as MP3 if listening at speed suits you better.
The transcript is in the language spoken, so a Mandarin drama produces Chinese text. There are 99 languages supported for transcription, but no translation step. If a link happens to carry a real English caption track, JotVox can return that track. Otherwise, take the Chinese transcript and translate it with the tool you prefer.
Yes. Any transcript exports as SRT or VTT with timings, which gives subtitlers a timed source-language base instead of typing from scratch. TXT is available when you only want the words, and MP3 when you want the audio alone. Timed exports work whether the text came from a caption track or from AI transcription.
AI transcription and AI summaries both run on credits, so longer content uses more than a short clip. New accounts get 3 free credits plus one free AI transcription per day after verifying an email address, which is enough to test a real episode. Credit packs start at $4.99 for 20 credits and subscriptions start at $9 per month.