The scrolling comments are not a script. JotVox writes down what the video actually says.
Get a Niconico transcript →Niconico (ニコニコ動画) is the site that made overlaid viewer comments a normal way to watch video, and that overlay is still the first thing anyone notices. It is worth being precise about what it is: a live-feeling layer of audience reactions, timed to moments in the video, written by viewers. It is not a transcript, not a caption track, and often not even a comment on what is being said.
Some Niconico videos also carry real text tracks, added by the uploader or as part of the video presentation. Where one is available, JotVox returns it in seconds, free, with no sign-up. Where there is none, which is common for commentary, gameplay and talk uploads, AI transcription works from the audio and produces the spoken content instead.
The distinction matters most for archival work. Save the comments and you have preserved an audience. Save the transcript and you have preserved the video's content.
The site's native content is some of the harder material to transcribe accurately:
JotVox's AI transcription is rated at 98%+ accuracy across 99 languages, and a clear-spoken Japanese talk upload transcribes very well. A layered game commentary with music and effects will produce more errors, and proper nouns from a niche fandom are the most likely thing to come back wrong. That is a known limit, not a surprise.
JotVox transcribes in the language spoken. A Japanese video produces Japanese text, across a set of 99 supported languages. It is not a translation tool, and it will not hand you an English version of a Japanese video.
Japanese text is still the unlock, because it is the format everything else accepts. You can put it through a translation service, read it with a dictionary tool, search it for a term you heard, or hand it to a language model for a summary. Going from fast spoken Japanese over game audio to accurate written Japanese is the difficult part, and that is the part JotVox does. See transcribing a video in any language for the general workflow.
Niconico transcripts get used for work that would be hard to do any other way:
Anyone doing this usually follows the same creators across sites: YouTube for mirrored uploads, Twitch for live archives, and Bilibili where Japanese content is reposted and discussed.
A handful of Niconico links come back empty, and the reasons are consistent:
If you already hold the file, upload it straight from your device, Google Drive or Dropbox. Exports come out as TXT, SRT, VTT or MP3, and the format guide covers which one you want.
1. Copy the video link Open the video in a browser and copy the address of the watch page, or use the share option in the app and copy the link it gives you.
2. Paste it into JotVox Paste it into the box on the JotVox home page. Archived files you already have can be uploaded directly from your device, Google Drive or Dropbox.
3. Take the captions or run AI transcription If the video has a real text track, JotVox returns it free and without an account. If not, AI transcription listens to the Japanese audio and writes the transcript.
No. Those comments are an audience layer written by viewers, not a record of the video's content, and JotVox returns only what is spoken in the audio. If you are archiving a video, the transcript preserves what the creator said. The comment overlay is a separate thing entirely and would need a different tool.
Japanese, if that is what is spoken. JotVox transcribes in the language of the audio and supports 99 languages, but it does not translate. To read a video in English, take the Japanese transcript and run it through a translation tool. Accurate written Japanese is a far better input for translation than the raw audio.
Usually yes, and often better than expected, because synthesised narration is evenly paced and clearly enunciated. Accuracy drops where the voice is heavily pitched or distorted for effect, or where it competes with game audio and music. As with any Japanese content, niche character names and coined subculture terms are the most likely words to come back wrong.
You can try, and results depend on the recording rather than its age. Low-bitrate audio from an old upload loses detail that speech recognition needs, so expect more errors than on a recent upload. It is still generally readable and searchable, which is usually the point for archival work where no better copy of the video exists.
Nothing, if the video carries an existing text track, since grabbing one is free with no sign-up. AI transcription uses credits, and most Niconico uploads need it. A new account gets 3 free credits plus one free AI transcription per day after email verification. After that, credit packs start at $4.99 for 20 credits and subscriptions start at $9 a month.