Download xhs videos
Download a Xiaohongshu (RedNote / 小红书) creator's videos to a local folder by driving a real logged-in Chrome. Tested on Claude Code (Claude in Chrome) and Codex (ChatGPT for Chrome) — each has its own route. Use when the user wants to batch-download / archive / 下载 / 抓取 a 小红书 (xiaohongshu / RedNote / xhs) 博主 / 用户 / up主 的视频 / 笔记, mirror a profile's video notes, or save someone's xhs videos for offline viewing. Triggers: 下载小红书视频, 抓小红书博主视频, 把这个博主的视频都下下来, download xiaohongshu videos, archive a RedNote creator. Personal/offline use only — not for re-posting or commercial use.From its SKILL.md
npx -y skills add YijiaDuan/download-xhs-videosAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 2 commands, including `scripts/download.sh <out-dir> "NN_<title>.mp4"` and 1 more.
SKILL.md
15.6 KB, ~4.0k tokens by cl100k_base, as published. Nobody here has run it
download-xhs-videos — batch-download a 小红书 creator's videos
Xiaohongshu gives you no download button (checked 2026-07: the web share panel offers only a QR code and "copy link", the … menu only "report", and the creator dashboard won't even export your own videos). The two obvious workarounds both fail: yt-dlp support is flaky, and a plain scraper hits signed-request (x-s/x-t) walls.
What works is driving a real, logged-in Chrome:
Real click on the note's TITLE to open it → read the video URL out of the page's
__INITIAL_STATE__→ download withcurl, outside the browser.
How the URL gets from the page to curl is the only part that differs per agent.
Supported agents
| Agent | Browser layer | Status | Route |
|---|---|---|---|
| Claude Code | Claude in Chrome extension | ✅ verified end-to-end 2026-07 | Route A |
| Codex (desktop / CLI) | chrome@openai-bundled, ChatGPT for Chrome | ✅ verified end-to-end 2026-07 | Route B |
| Anything else | — | ❓ not tested | — |
Only these two have been tested. Other coding agents may have the needed pieces — a trusted click into the real logged-in Chrome, a way to read page JS state, and a shell for curl — but nobody has run this skill on them. If you are a different agent: read The invariants, map them onto your own browser API, and tell the user you are improvising off an untested path.
Before you start — preconditions
- macOS. (Route A needs
pbpaste; the extension hosts are mac-only today.) - Your browser layer is connected to a Chrome that is logged into 小红书.
- The creator's profile URL —
https://www.xiaohongshu.com/user/profile/<uid>. A bare uid URL works when you're logged in; an?xsec_token=…copy works too. - A target folder, default
~/Documents/xhs-<handle>/.
Cross-account works. The logged-in account does not have to be the creator you're downloading — verified on both routes with a profile belonging to someone other than the logged-in user. You just need some valid 小红书 session so pages render.
Ethics gate — do this, don't skip it
Open the profile and read the creator's bio first. Many creators write 「原创作品,禁止搬运和商用」. If so, confirm with the user that this is personal/offline viewing only. Downloading public videos for yourself is a defensible grey area; re-publishing or monetizing someone's flagged original work is not — refuse that. State what the bio says and get a clear yes before mass-downloading.
The invariants
Properties of Xiaohongshu and Chrome, not of any agent. They hold on every route.
-
A note only opens on a real click.
location.href, a hand-built<a>, or navigating to a note URL all get bounced back to/exploreby XHS's router. Only a genuine mouse click at the card's coordinates loads the note with its video stream. -
Click the card's TITLE, not its cover. Each card has three anchors to the same note: a zero-size one,
a.cover(223×297), anda.title. Clicking the cover's center is swallowed by XHS's 「图搜同款」 visual-search hover overlay — the click lands, nothing navigates, and you burn a loop iteration wondering why.a.titlehas no overlay. -
Wait for the cover images before trusting any rect. Until the covers finish loading, the masonry layout has not run and every card returns the same
getBoundingClientRect()(observed: all five first cards reportingx=307). Clicking that opens the wrong note or nothing. The tell is duplicate centers across cards. Fix: scroll down ~5 ticks, scroll back up, wait ~2 s — that forces the lazy-load. Once loaded, centers are correctly distinct (307 / 562 / 817 / 1072 / 1327).scripts/02-locate-card.jsrefuses to guess when it detects duplicates. -
The stream's codec keys are not stable — never hardcode them. They used to be
h264 / h265 / av1 / h266. As of 2026-07 live pages shipEF4 / EF5 / EF6 / EF7(EF4 is their h264 ladder — a downloaded EF4 stream probes ash264). Worse: key order is not preference order (observedEF4, EF6, EF5, EF7) and some keys are empty arrays. Walk a preference list, then fall back to whatever keys exist, skip empties, and never index blindly. (This bug silently broke the skill for everyone until it was fixed.) -
Don't
fetch()the video inside the page. Chrome caps ~6 connections per host; a few hung fetches to the CDN exhaust the pool and every later fetch hangs forever.curlsidesteps this. -
The CDN URL is
http://, nothttps://, and needs a Referer. Pass--referer https://www.xiaohongshu.com/plus a normal User-Agent. (Use curl's--refererflag rather than spelling the header out with-H— Xiaohongshu's own upload gateway runs an Aliyun WAF that false-positivescurl -H "…"as header injection and 405s the request. Same for||chains; see Verify.) -
Only one agent per tab. If two browser-driving agents both attach to the same tab, you get endless CDP timeouts and kernel resets that look exactly like a missing capability. Claim a tab nobody else is on.
-
Heavy automated use trips XHS's security verification. After enough automated navigation in one session, profile loads start redirecting to
/website-login/captcha?…verifyType=…and the page has no feed at all (noteAnchors: 0). This is rate-limiting, not a bug in your code. Do not attempt to solve the CAPTCHA. Stop, tell the user their session hit a verification wall, and ask them to clear it themselves in that Chrome window. Then resume. Pace the loop (the per-note navigation is already slow enough for normal use) and don't re-run the whole profile repeatedly while debugging — test on 1–2 notes. -
Quality:
masterUrlis the top stream of the chosen ladder; XHS source bitrate is modest, so 1–5 MB for a 2-minute clip is normal, not a bug.
Route A — Claude Code
Browser layer: the Claude in Chrome extension (list_connected_browsers, navigate, tabs_context_mcp, javascript_tool, computer, browser_batch).
Two extra traps specific to this route
-
Chrome silently blocks automation-triggered downloads. A blob download via
a.click()from injected JS needs a user-activation gesture that injected code doesn't carry — and the extension's synthetic clicks don't grant download activation either (they fire DOMclickhandlers, soexecCommand('copy')works, but the download is dropped with no error and no file). So: never download through the browser. Usecurl. -
The harness blocks query-string data from your context. If in-page JS returns the video URL, the tool result is
[BLOCKED: Cookie/query string data]because of thexsec_token. So you can't read the URL and build acurlcommand with it. Route the URL clipboard →curlso it never touches your context. -
javascript_tooldoes not await an async IIFE. Anasync () => {…}returns a pending Promise and you receive{}— the side effects still happen, but the return value is lost. All three scripts here are therefore synchronous (except part 1 of01, whose return value genuinely doesn't matter). If you write your own, keep them sync or you'll never get your coordinates back.
Recipe
Step 0 — Connect and open the profile. Connect the browser, open a fresh tab (tabs_context_mcp createIfEmpty:true), navigate to the profile URL, get_page_text to read the bio (ethics gate) and the note titles.
Step 1 — Load every note and count. Run part 1 of scripts/01-load-and-count.js (async scroll, return value ignored), then part 2 (sync) to read the count. Re-run part 2 until growing is false. Report the total to the user — it counts image + video notes both.
Step 2 — Per-note loop. For note index N (0-based, DOM order):
- Call A — locate (one
browser_batch):navigateback to the profile, thenscripts/02-locate-card.jswithNset. It returns{count, x, y, text, ready}in screenshot pixels. If it returnserr: 'masonry-not-laid-out', scroll down 5 / up 10, wait 2 s, and re-run. - Call B — open + extract + copy (one
browser_batch):computer left_clickthe returned coords → the note opens (you're clicking the title).- Wait ~5 s, then
javascript_toolrunsscripts/03-extract-and-copy.js. It returns{ok, codec, height, urlLooksRight, buttonPlaced}. Onerr: 'no-video-yet', wait and re-run; onerr: 'no-video'it's an image note — skip it. computer left_click [150,155]— a trusted click on the copy button → URL is on the clipboard.
- Call C — download (Bash):
scripts/download.sh <out-dir> "NN_<title>.mp4".
Step 3 — Verify. See Verify.
Route A gotchas
- Viewport vs screenshot pixels differ ~2%. Scale rect coords by
1496/innerWidthand812/innerHeight(the locate script does this) or clicks drift on the rightmost column. backnavigation re-renders the DOM, so prior refs are stale. Rebuild the card list every loop.- Browser downloads vanish both ways: "Ask where to save each file" on → a native Save dialog the extension can't touch; off → no user activation. Stop fighting it.
Route B — Codex
Browser layer: the chrome@openai-bundled plugin driven through mcp__node_repl__js. Check codex plugin list shows it installed, enabled, and that ChatGPT for Chrome is installed in the Chrome profile that's logged into 小红书.
Route B is simpler than Route A — it has neither of Route A's first two traps. There is no context guardrail on token-bearing URLs, so you can just read the URL and curl it. Do not port the clipboard trick over: you don't need it, and it cannot work here.
Setup
// inside mcp__node_repl__js, once per session
if (globalThis.agent?.browsers == null) {
const { setupBrowserRuntime } = await import(
"<codexHome>/plugins/cache/openai-bundled/chrome/<version>/scripts/browser-client.mjs");
await setupBrowserRuntime({ globals: globalThis });
}
await chrome.nameSession("📥 下载小红书视频");
const tabs = await chrome.user.openTabs();
const tab = await chrome.user.claimTab(tabs[i]); // claim a tab NO ONE ELSE is on
Read await chrome.documentation() once — it's the authoritative spec and longer than this summary.
Two hard constraints on this route
-
tab.playwright.evaluateruns in a READ-ONLY scope. Global assignment and DOM mutation both throw:Error: global assignment is not available in playwright.evaluate because the DOM is read-onlySo Route A's "stash the URL on
window.__urlForCopyand inject a copy button" is impossible. Fine — reading values andreturning them is allowed, which is all you need. -
window.__INITIAL_STATE__is NOT visible from that scope. Reading it directly returns an empty object (noteDetailMapkeys[],hasStream:false) — the evaluate world is isolated from the page's realwindow. You must parse the inline<script>instead, and replace bareundefinedwithnullbeforeJSON.parse(XHS serializes literalundefined, which is not valid JSON):const raw = Array.from(document.scripts) .find(el => (el.textContent || '').includes('window.__INITIAL_STATE__='))?.textContent || ''; const prefix = 'window.__INITIAL_STATE__='; const state = JSON.parse( raw.slice(raw.indexOf(prefix) + prefix.length).replace(/;\s*$/, '').replace(/\bundefined\b/g, 'null'));
Recipe
Step 0 — claim a fresh tab, goto the profile, read the bio (ethics gate) via tab.playwright.domSnapshot().
Step 1 — scroll to the bottom, count distinct note ids. Port 01's logic as pure, value-returning evaluate calls (no window assignments).
Step 2 — per-note loop:
- Locate —
evaluatethat returnsa.title's center for cardN. Apply invariants 2 and 3 (title anchor; refuse duplicate rects until covers load). - Open —
await tab.cua.click({x, y, button: 1}). Real trusted click; verified to open the note. Confirm withawait tab.url(). - Extract —
evaluatethat parses the inline state per constraint 2 above, picks the stream per invariant 4, and returns the URL string:const stream = state?.note?.noteDetailMap?.[noteId]?.note?.video?.media?.stream; for (const k of ['h264','h265','av1','h266','EF4','EF5','EF6','EF7']) { const it = stream?.[k]?.[0]; if (it) return it.masterUrl || it.backupUrls?.[0] || null; } - Download — plain shell, no clipboard:
curl -sS -L -A "Mozilla/5.0" --referer https://www.xiaohongshu.com/ \ -o "NN_<title>.mp4" "$URL" - Back to the profile, repeat. Rebuild the card list each loop.
Step 3 — Verify. See Verify.
Route B gotchas
- CDP flakiness is normal.
Runtime.evaluatetimes out and the JS kernel resets (js execution timed out; kernel reset, rerun your request), sometimes several times running. Usetimeout_ms: 60000, re-runsetupBrowserRuntime()whenglobalThis.agent?.browsers == null, re-claimTab, retry the same step. Do not conclude the capability is missing and switch methods — the retry works. (If it never recovers, check invariant 7: another agent on your tab.) - Codex has extras this skill doesn't use:
tab.capabilities.get("pageAssets")→list()/bundle()can inventorykind:"video"assets and export them to a local directory;downloadMedia({x,y})triggers a real media download andpath()returns the file;clipboard.readText()exists too. These are plausible shortcuts — untested here, so treat them as experiments, not the documented path.
Verify
After the loop: count matches, none <10 KB, all ISO Media.
cd <out-dir>
for f in *.mp4; do
if ! file "$f" | grep -q 'ISO Media'; then echo "NOT_MP4: $f"; fi
done
ls -1 *.mp4 | wc -l; du -sh .
If ffprobe is available, spot-check that durations match what the profile showed — that's how you catch truncated downloads.
Files in this skill
SKILL.md— this playbook.scripts/01-load-and-count.js— two parts: async scroll-to-bottom, then a sync counter.scripts/02-locate-card.js— sync; returns cardN's title-anchor center, with the masonry sanity check.scripts/03-extract-and-copy.js— sync; codec-agnostic stream pick + the clipboard copy button. Route A only (the button can't exist on Route B).scripts/download.sh— clipboard →curlwith a CDN guard. Route A only; Route B curls the URL it already has.
Dependencies (explicit)
- macOS (
curl;pbpastefor Route A). - A Chrome signed into 小红书, plus one of:
- Claude Code with the Claude in Chrome extension connected, or
- Codex with
chrome@openai-bundledenabled and ChatGPT for Chrome installed in that profile.
Related skills
waitlist-farmer— same "drive a real logged-in browser through a repetitive flow" muscle, different domain.
What ships with it: 8 files
38.6 KB alongside SKILL.md, 4 of them executable
scripts/
- 01-load-and-count.jsruns2.3 KB
- 02-locate-card.jsruns3.3 KB
- 03-extract-and-copy.jsruns3.9 KB
- download.shruns1.4 KB
- .gitignore293 B
- LICENSE1.0 KB
- README_EN.md12.9 KB
- README.md13.5 KB