JJotVox Open the app →
Bilibili transcripts

Bilibili video to text, without scraping Danmaku

Paste a Bilibili link and get the spoken words as a clean transcript, whether the video carries a subtitle track or nothing at all.

Get a Bilibili transcript →

Does a Bilibili video have subtitles, or is that just Danmaku?

Danmaku (弹幕) is the stream of viewer comments that flies across the picture. It is a comment layer, not a caption track. It reacts to the video, repeats running jokes, and frequently has nothing to do with the sentence the speaker is saying at that moment. Collecting Danmaku gives you an audience reaction log. It does not give you the content of the video.

Bilibili (哔哩哔哩) does carry real subtitle tracks on part of its catalogue. Some come from the uploader, some come from viewers, since the site has long supported people contributing subtitles to videos they care about. Coverage is uneven and unpredictable. A popular channel's tutorial may have a tidy Chinese track and sometimes an English one alongside it, while a 40-minute lecture upload right next to it has no text of any kind.

JotVox checks for a genuine caption track first. If one is there, you get it in seconds, free, with no sign-up. If there is none, AI transcription listens to the audio and writes the transcript itself.

Why Bilibili audio is harder than a talking-head clip

The kind of video people want transcribed from Bilibili tends to be the kind that punishes lazy speech recognition:

JotVox's AI transcription is rated at 98%+ accuracy across 99 languages, and clean studio-style narration lands near the top of that. A heavily edited compilation with music under every line will come back lower. We wrote about what that number does and does not mean in our note on transcription accuracy.

Can I read a Bilibili video in English?

Here is the honest version. JotVox transcribes the language that is actually spoken. A Mandarin video comes back as Chinese text, a Japanese one as Japanese text, and so on across 99 supported languages. It is a transcription tool, not a dubbing or translation service.

That still solves most of the problem. Once the speech is text you can paste it into whatever translation tool you already trust, read it at your own pace, search it for the one term you came for, or feed it to a language model. Going from spoken Mandarin to readable Chinese text is the step that was blocking you. There is more on the workflow in reading a Bilibili video in English and in Bilibili video to text in Chinese.

What people pull Bilibili transcripts for

Bilibili's centre of gravity is study material. A large share of the requests we see are people trying to extract something teachable:

If you track the same creators across platforms, the same paste-a-link flow works on YouTube, on Niconico for the Japanese side of that world, and on Xiaohongshu for the short vertical version.

When a Bilibili link comes back empty

Not every link is reachable, and it is better to know why than to keep retrying:

For the wider version of this problem, see getting text out of a video with no captions.

Exports, timestamps and summaries

Every finished transcript can leave as TXT for reading and pasting, SRT or VTT if you need timed subtitle lines, or MP3 if you want the audio on its own. The format comparison covers which to pick.

For hour-long lectures there are AI summaries too: key takeaways, chapters, pull quotes and a timeline, so you can decide whether the full text is worth reading before you read it. Summaries and AI transcription both use credits. Grabbing an existing subtitle track does not.

How to transcribe a Bilibili video

    1. Copy the video link Open the video in a browser and copy the address from the address bar, or use the share option in the Bilibili app and copy the shortened link it offers. Both forms work.

    2. Paste it into JotVox Drop the link into the box on the JotVox home page. You can also upload a file straight from your device, Google Drive or Dropbox if you already have the video saved.

    3. Let JotVox pick the method If the video has a real subtitle track, JotVox returns it right away, free and without an account. If it has none, run AI transcription on the audio instead.

    4. Export or summarize Download the transcript as TXT, SRT, VTT or MP3, or generate a summary with takeaways, chapters and quotes for longer lectures and reviews.

Frequently asked questions

Can JotVox pull the Danmaku comments from a Bilibili video?

No. Danmaku is a layer of viewer comments drawn over the video, and JotVox returns only what is spoken in the audio. That is usually what people actually want. The comments tell you how an audience reacted to a moment, while the transcript tells you what the creator said, which is the part you can study, quote or search.

What language will my Bilibili transcript be in?

The language spoken in the video. A Mandarin video produces Chinese text, and JotVox supports 99 languages for AI transcription. It does not translate. If you need English, run the finished transcript through the translation tool you prefer. The hard part, turning fast spoken Mandarin into accurate written text, is the step JotVox handles.

Is transcribing a Bilibili video free?

Grabbing an existing subtitle track is free and needs no sign-up. AI transcription, used when a video has no caption track, runs on credits. New accounts get 3 free credits plus one free AI transcription per day once the email is verified. Beyond that, credit packs start at $4.99 for 20 credits and subscriptions start at $9 per month.

Can I get subtitle files for a Bilibili video, not just plain text?

Yes. Any transcript can be exported as SRT or VTT with timings, as TXT if you only want the words, or as MP3 if you want the audio itself. SRT suits most video editors and players, while VTT is the web-native option. Timed exports are available whether the text came from an existing caption track or from AI transcription.

How long does a one-hour Bilibili lecture take?

If a subtitle track exists, seconds. AI transcription of a full hour takes longer because the audio has to be processed end to end, but it runs in the background and you do not need to sit on the page. For long lectures, generating chapters afterwards is usually faster than skimming the raw text yourself.

Keep reading

Browse all supported platforms.