JJotVox Open the app →
Xiaohongshu / RedNote

Turn a Xiaohongshu post into text you can search

Short vertical video, spoken fast over music, usually with no caption track at all. That is the case JotVox is built for.

Get a Xiaohongshu transcript →

Do Xiaohongshu videos come with captions?

Usually not in any form a tool can read. Xiaohongshu (小红书), known internationally as RedNote, is built around short vertical posts, and creators there rarely publish a separate subtitle track. What looks like subtitles is normally text burned into the picture by the editing app: the caption a creator typed onto the clip, a price tag, a product name, a punchline sitting across the bottom third.

That distinction matters for what you get back. Burned-in text is part of the image, not a text layer, so it does not appear in a transcript. JotVox transcribes the spoken audio. When a creator says the brand name out loud, it lands in the transcript. When they only put it on screen and never say it, it does not.

In practice that is still where the useful information lives on Xiaohongshu, because the honest opinion is almost always spoken, not typed on the overlay.

Why short vertical video is a hard case for speech recognition

A 45-second Xiaohongshu clip packs more difficulty per second than an hour-long lecture:

JotVox's AI transcription is rated at 98%+ accuracy and covers 99 languages, and it handles clear speech over music well. A whispered aside under a loud track is where any system loses ground. If a post genuinely has no usable audio, no tool can invent one.

Product research, trend work and what a review actually claimed

Xiaohongshu is a search engine for buying decisions, which makes text out of it commercially useful:

Teams doing this work rarely stop at one platform. The same link-paste flow covers TikTok for the international short-form equivalent, Weibo for the discussion around a brand, and Bilibili for the long-form review versions of the same products.

What language does the transcript come out in?

The one being spoken. A Mandarin post gives you Chinese text; JotVox supports 99 languages for AI transcription and writes what it hears. It does not translate for you, and we would rather say that plainly than let you find out after paying.

For most research work that is enough. Chinese text can be pasted into any translation tool, skimmed with a browser translator, searched for a brand name, or handed to an assistant to summarize. See Xiaohongshu video to text in Chinese for the full walkthrough, or transcribing a video in any language for the general case.

Posts that will not transcribe

A few Xiaohongshu links come back with nothing, and the reason is almost always one of these:

If you already have the file on your phone or in cloud storage, uploading it directly sidesteps all of this. JotVox accepts uploads from your device, Google Drive and Dropbox.

How to transcribe a Xiaohongshu (RedNote) video

    1. Copy the post link Open the video post, use the share option and copy the link it gives you. On desktop you can copy the address of the post page directly instead.

    2. Paste it into JotVox Put the link in the box on the JotVox home page. If you saved the clip already, upload the file from your device, Google Drive or Dropbox instead.

    3. Transcribe the spoken audio Most Xiaohongshu posts have no caption track, so AI transcription runs on the audio and returns the words the creator actually said, in the language they said them.

    4. Export it into your research Take the text as TXT for notes and docs, SRT or VTT if you need timings, or MP3 for the audio. Add an AI summary to pull the key claims out of a batch of reviews.

Frequently asked questions

Will the text overlaid on a Xiaohongshu video appear in my transcript?

No. Text a creator burns into the frame with an editing app is part of the picture, not a caption track, so it will not show up. JotVox transcribes the spoken audio. In practice the substance of a review is spoken rather than typed on screen, so the transcript usually captures the claims you were looking for anyway.

Can I get a Xiaohongshu video in English?

The transcript comes back in the spoken language, so a Mandarin post produces Chinese text. JotVox transcribes across 99 languages but does not translate. Once you have accurate Chinese text you can run it through any translation tool, which is far more reliable than trying to translate straight from fast spoken audio over a music bed.

Do I need an account to transcribe a Xiaohongshu post?

Not if the video happens to carry a caption track, since grabbing an existing track is free with no sign-up. Xiaohongshu posts rarely do, so most will need AI transcription, which uses credits. Verifying your email gets you 3 free credits plus one free AI transcription a day. Credit packs start at $4.99 for 20 credits.

Does background music ruin the transcript?

It lowers accuracy rather than blocking transcription. Xiaohongshu clips almost always run music under the voice, and AI transcription handles that in most cases, rated at 98%+ accuracy on clear speech. Quiet asides mixed under a loud track are where errors show up. Louder, closer voiceover recorded on a phone in a quiet room transcribes very cleanly.

What happens to my audio after transcription?

Audio is processed in memory and deleted right after the transcription finishes. Nothing is kept sitting in storage waiting to be dealt with later. That matters if you are running competitor research or client work through the tool and would rather not leave copies of source material anywhere you cannot see them.

Keep reading

Browse all supported platforms.