Youtube transcribe skill
Skill feiskyer/claude-code-settings/skills/youtube-transcribe-skill
Extract subtitles/transcripts from YouTube videos. Triggers: "youtube transcript", "extract subtitles", "video captions", "视频字幕", "字幕提取", "YouTube转文字", "提取字幕".From its SKILL.md
npx -y skills add feiskyer/claude-code-settings --skill youtube-transcribe-skillAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- reads credentialsReads from 1 credential source: `--cookies-from-browser=chrome`.
- runs commandsInstructs the agent to run 3 commands, including `which yt-dlp` and 2 more.
SKILL.md
4.5 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
YouTube Transcript Extraction
Extract subtitles/transcripts from a YouTube video URL and save them as a local file.
Input YouTube URL: $ARGUMENTS
Step 1: Verify URL
Confirm the input is a valid YouTube URL (supports youtube.com/watch?v=, youtu.be/, and youtube.com/shorts/ formats). If no URL is provided via arguments, check the conversation context for a YouTube link.
Step 2: CLI Quick Extraction (Priority Attempt)
Use command-line tools to quickly extract subtitles.
2.1 Check Tool Availability
Execute which yt-dlp.
- If
yt-dlpis found, proceed to 2.2. - If
yt-dlpis not found, skip to Step 3.
2.2 Get Video Title
yt-dlp --cookies-from-browser=chrome --get-title "[VIDEO_URL]"
- Tip: Always add
--cookies-from-browserto avoid sign-in restrictions. Default tochrome. - If it fails with a browser error (e.g., "Could not open Chrome"), ask the user to specify their available browser (e.g.,
firefox,safari,edge) and retry.
2.3 Download Subtitles
yt-dlp --cookies-from-browser=chrome --write-auto-sub --write-sub --sub-lang zh-Hans,zh-Hant,en --skip-download --output "<Video Title>.%(ext)s" "[VIDEO_URL]"
2.4 Convert to Plain Text
yt-dlp saves subtitles as .vtt or .srt files. Convert the downloaded file to plain Timestamp Text format:
- Read the downloaded subtitle file (
.vttor.srt). - Strip VTT/SRT headers, styling tags, and duplicate lines.
- Save as
<Video Title>.txtwith oneTimestamp Textentry per line.
2.5 Verify Results
- Exit code 0: Convert and save the subtitle file, then report completion.
- Exit code non-0:
- If error is related to browser/cookies, ask user for correct browser and retry.
- If other errors (e.g., video unavailable), proceed to Step 3.
Step 3: Browser Automation (Fallback)
When the CLI method fails or yt-dlp is missing, use Chrome DevTools MCP to extract subtitles via browser UI automation.
3.1 Check Tool Availability
Check if Chrome DevTools MCP tools are available (look for tools matching chrome__new_page or similar).
If Chrome DevTools MCP is not available and yt-dlp was not found in Step 2, stop and notify the user: "Unable to proceed. Please either install yt-dlp (for fast CLI extraction) or configure Chrome DevTools MCP (for browser automation)."
3.2 Open Video Page
Use Chrome DevTools MCP new_page to open the video URL.
3.3 Analyze Page State
Use Chrome DevTools MCP take_snapshot to read the page accessibility tree.
3.4 Expand Video Description
The "Show transcript" button is usually hidden within the collapsed description area.
- Search the snapshot for a button labeled "...more", "...更多", or "Show more" (in the description block below the video title).
- Use Chrome DevTools MCP
clickto click that button.
3.5 Open Transcript Panel
- Use Chrome DevTools MCP
take_snapshotto get the updated UI. - Search for a button labeled "Show transcript", "显示转录稿", or "内容转文字".
- Use Chrome DevTools MCP
clickto click that button. - If the button is not found, the video may not have a transcript available — notify the user and stop.
3.6 Extract Content via DOM
Directly reading the accessibility tree for long transcript lists is slow and token-heavy. Use Chrome DevTools MCP evaluate_script to run this JavaScript instead:
() => {
const segments = document.querySelectorAll("ytd-transcript-segment-renderer");
if (!segments.length) return "BUFFERING";
return Array.from(segments)
.map((seg) => {
const time = seg.querySelector(".segment-timestamp")?.innerText.trim();
const text = seg.querySelector(".segment-text")?.innerText.trim();
return `${time} ${text}`;
})
.join("\n");
};
If it returns "BUFFERING", wait a few seconds and retry (up to 3 attempts).
3.7 Save and Cleanup
- Save the extracted text as
<Video Title>.txt. - Use Chrome DevTools MCP
close_pageto release resources.
Output Requirements
- Save the subtitle file to the current working directory.
- Filename format:
<Video Title>.txt - File content format: Each line should be
Timestamp Subtitle Text. - Report upon completion: file path, subtitle language, and total number of lines.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most video audio skills give in ~1.1k tokens
Counted across 619 of the 725 authors here whose files we hold, read 2026-09-06
- Read product marketing context firstin 13 of 619, across 7 files
- Define the core visual thesis in one sentencein 11 of 619, across 3 files
- Break the concept into 3 to 6 scenesin 11 of 619, across 3 files
- Render the smallest working version firstin 11 of 619, across 3 files
- Start with a low-quality smoke test renderin 11 of 619, across 3 files
- Add captions for accessibility and engagementin 11 of 619, across 5 files
- Write the scene outline before writing codein 11 of 619, across 3 files
- Specify subject, action, camera, style, and moodin 11 of 619, across 5 files
- Decide what each scene provesin 10 of 619, across 2 files
- Export one clean thumbnail framein 10 of 619, across 2 files
- Pick the right tool for the jobin 10 of 619, across 4 files
- Run the test suite before proposing a fixin 8 of 619, across 7 files
Said here and by no other author read
- Verify the input is a valid YouTube URL
- Check tool availability using which yt-dlp
- Open the video page using Chrome DevTools MCP
- Expand the video description
- Open the transcript panel
- Extract transcript content via DOM
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.