agentsclimarketplace

Transcript fetcher

Skill MatrixFounder/Universal-skills/skills/transcript-fetcher

Collection of high-leverage "Meta-Skills" designed to upgrade AI Agents from simple chat bots to autonomous engineers

Install
npx -y skills add MatrixFounder/Universal-skills --skill transcript-fetcher

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when the user wants a clean plain-text transcript of a video URL (YouTube, Vimeo, X.com/Twitter incl. Broadcasts/Spaces) or a Skool classroom lesson. X uses embedded captions when present and falls back to ASR (MacWhisper/Whisper/whisper.cpp/cloud) automatically. Skool cookies are optional (public communities work without; private ones accept Netscape cookies.txt). Manual->auto language fallback, rolling caption dedup, >> speaker turns preserved, JSON stat sidecar plus optional .description.md sidecar.

SKILL.md

28.4 KB, as published. Nobody here has run it

Transcript Fetcher

Purpose: Given a video URL, produce a clean plain-text transcript ready for downstream summarization or analysis. Single responsibility: fetch + clean. Composes well with summarizing-meetings (run this first, then feed the resulting .txt to that skill).

Dependencies: yt-dlp is vendored in scripts/.venv (installed by scripts/install.sh) — it is NOT on $PATH and never will be. Do NOT which yt-dlp, pip install yt-dlp, or brew install yt-dlp globally — a missing $PATH entry does not mean yt-dlp is absent, it means the check was wrong; this exact false-negative once caused a manual out-of-skill workaround. The two canonical readiness probes are scripts/.venv/bin/python -m yt_dlp --version and scripts/.venv/bin/python scripts/fetch.py doctor (§4 Script Contract; the latter also reports ffmpeg + ASR backends — see the ASR portability note under §5 Safety Boundaries). Callers that shell in (e.g. a downstream integrator) MUST invoke the venv interpreter directly (scripts/.venv/bin/python), never a $PATH python.

1. Red Flags (Anti-Rationalization)

STOP and READ THIS if you are thinking:

  • "I'll just paste the URL into the model and ask it to transcribe" -> WRONG. Models do not have audio access and will hallucinate. Always run fetch.py and read back the resulting .txt.
  • "Manual ru failed, I'll just use en auto-translation, no warning needed" -> WRONG. Auto-translated English subtitles destroy idioms, names, and technical terms. The quality_flag: english_auto_translation in the stat MUST be surfaced to the user.
  • "I'll skip writing the JSON stat sidecar, the .txt is enough" -> WRONG. The sidecar records WHICH track was picked. Without it, downstream cannot tell whether the transcript is high-quality manual subs or low-quality auto-translation.
  • "The transcript has weird >> markers, I'll strip them" -> WRONG. Those are speaker-turn boundaries. Removing them collapses multi-speaker meetings into a single voice and ruins downstream attribution.
  • "Auto-generated Russian (ru-orig) is garbage, I'll prefer en instead" -> WRONG. ru-orig is the actual Russian audio transcribed; ru (without -orig) is often an English auto-translation back to Russian. ru-orig > ru > en.
  • "I'll add ffmpeg to the pip deps just in case" -> WRONG. The caption path (WebVTT parsing) is pure Python and needs nothing extra; ffmpeg is a soft-optional external system tool, never a pip dependency. It IS genuinely required for the X ASR path on HLS sources (Broadcasts/Spaces) — yt-dlp uses it to extract a clean audio-only m4a, and the skill fails fast (exit 7) when it is absent there — but it is detected at runtime (install_components.py), not bundled. Do not pull heavy packages into requirements.txt.

2. Capabilities

  • Fetch YouTube and Vimeo captions via yt-dlp (no audio download).
  • Fetch X.com / Twitter — native status video AND Broadcasts/Spaces. The X provider is captions-first: it reuses embedded subtitles/automatic_captions when present, and only when none exist does it download the smallest media and transcribe via ASR. Fully automatic — no mode switch. ASR runs through a pluggable backend chain: MacWhisper (mw) → Whisper CLI → whisper.cpp → opt-in OpenAI/compatible cloud. ffmpeg is required for the X ASR path on HLS sources (Broadcasts/Spaces): with it the smallest media is extracted to a clean audio-only m4a; without it the skill fails fast (exit 7) because yt-dlp's no-ffmpeg HLS output is not a valid container the ASR engine can open. (For non-HLS progressive media, or when embedded captions exist, ffmpeg is not needed.)
  • Fetch Skool lesson pages via a stdlib HTML scrape — public communities work without auth; private/paid ones accept an optional Netscape cookies.txt. Then delegate embedded YouTube/Vimeo videos to those adapters; capture author-supplied transcript field when present.
  • Fall back through a configurable ladder (default for ru: manual ru -> auto ru-orig -> auto ru -> auto en).
  • Clean captions to plain text — WebVTT, and for X also SRT and TTML/DFXP (vtt/srt/ttml/best preference list, so a non-VTT track is no longer skipped to ASR): strip timestamps, inline timing tags, and rolling-caption overlap; decode HTML entities. TTML is parsed safely (DTD/entity declarations refused — XXE/billion-laughs guard). For X, captions-first is language-robust: if the requested --lang has no track but the post carries captions in another language, those are used (manual preferred, with a note) rather than dropping to ASR.
  • Preserve >> speaker-turn markers as paragraph breaks.
  • Emit a JSON stat sidecar (chosen track, char count, speaker-turn count, quality flag, plus optional title/uploader/duration metadata). For X media it also records transcript_origin (embedded-captions | macwhisper | whisper-cli | whisper-cpp | openai-api) so downstream skills know HOW the text was produced.
  • Optionally write <out>.description.md (YAML frontmatter + Markdown body) when --with-description is passed — gives you a ready-to-ingest description for RAG / Obsidian / human review.
  • Batch mode for processing multiple URLs from a text file.
  • Source-agnostic architecture: each platform is one file under scripts/sources/. The yt-dlp + ASR pipeline is shared (sources/_ytdlp_media.py + asr/), so a future TikTok/Twitch/Vimeo-ASR provider is one new file + one host entry — no pipeline changes. Zoom and podcast slots remain reserved.
  • Configurable + secrets-safe: a skill-local .env (see scripts/.env.example) externalises every endpoint, model, and tool path. The cloud ASR endpoint works with any OpenAI-compatible server (Groq, self-hosted whisper). A .env holding an API key is refused unless chmod 600 (and not a symlink); the key is sent only in an HTTP header, never on argv or in logs.

3. Execution Mode

  • Mode: script-first
  • Why this mode: Fetching captions, parsing WebVTT, deduplicating rolling captions, and applying the fallback ladder are deterministic operations with > 5 lines of business logic. The CLI is the contract; SKILL.md is orchestration.

4. Script Contract

  • Install (one-time): creates the venv + yt-dlp, then reports which optional ASR components are present:
    bash skills/transcript-fetcher/scripts/install.sh
    # Optional ASR engines (for caption-less X media) — detect / install:
    ./scripts/.venv/bin/python scripts/install_components.py            # status report
    ./scripts/.venv/bin/python scripts/install_components.py --install-whisper   # pip openai-whisper into the venv
    ./scripts/.venv/bin/python scripts/install_components.py --system --run      # brew/apt ffmpeg + whisper.cpp
    
  • Readiness check (doctor, no network, import-free): answers "is this skill ready?" without $PATH guessing — the canonical replacement for which yt-dlp:
    ./scripts/.venv/bin/python scripts/fetch.py doctor           # human-readable report
    ./scripts/.venv/bin/python scripts/fetch.py doctor --json    # {v, interpreter, in_venv, ready, components, remediation}
    
    Reports the resolved interpreter, in-venv flag, yt-dlp version, ffmpeg, each ASR backend, and cloud opt-in state (key presence only, never the key). remediation names a flow-blocking gap for EVERY such gap — yt-dlp missing, ffmpeg missing (it back-stops X Broadcast/Space HLS ASR and whisper/whisper.cpp at runtime even though it is itself optional), or no ASR capability at all (no local backend AND no fully-configured cloud) — but never an individual missing ALTERNATIVE local ASR engine while ASR capability already resolves elsewhere (mw present, another local engine present, or cloud configured); those show only as informational install-hint lines in the human report. A fully-configured cloud backend (--asr-allow-cloud + key) genuinely suppresses the no-local-ASR hint — it never appears in remediation, and the human report gets one informational note instead of a demand to install a local engine it does not need. Exit 0 when yt-dlp is present (regardless of remediation), 7 when it is not.
  • Single URL:
    cd skills/transcript-fetcher
    ./scripts/.venv/bin/python scripts/fetch.py <URL> --out <path/to/output.txt>
    
    Optional flags: --lang ru (default), --prefer manual|auto (default manual), --with-description, --description-only, --cookies-file PATH, --json-errors, --debug (stage logging to stderr), --asr-allow-cloud (opt-in cloud ASR), --asr-model <id>, --asr-timeout-sec N, --max-duration-min N (X: transcribe only the first N minutes — clips the download when ffmpeg is present, the only case where the download itself is clipped; also clips the media-timeout floor below in that case), --keep-silence (X ASR: do NOT strip long silences before transcription — silence removal is ON by default to cut Whisper hallucinated filler on silent lead-in/out), --auth-map PATH (per-host cookies), --cookies-from-browser BROWSER (X: load cookies from a local browser via yt-dlp), --concurrent-fragments N (X: parallel HLS fragment downloads for the media, default 8; 1 = serial; values above 32 are capped at 32; CLI rejects <= 0 with exit 2; env TRANSCRIPT_FETCHER_CONCURRENT_FRAGMENTS non-positive/malformed falls back to 8 — these are three DISTINCT layers, not one shared clamp), --media-timeout-sec N (X: per-attempt budget for the media download only, separate from the probe's --timeout-sec; default duration-derived min(21600, max(600, duration*4))s — capped at 6h, and derived from the --max-duration-min-clipped duration when that flag is set AND ffmpeg is present (the only case where the download itself is clipped) — else 1800s).
  • X.com / Twitter (captions-first, automatic ASR fallback; no mode switch). A Broadcast/Space usually has no captions → ASR via the first available local backend (MacWhisper, etc.):
    cd skills/transcript-fetcher
    ./scripts/.venv/bin/python scripts/fetch.py \
        "https://x.com/i/broadcasts/<id>" \
        --out broadcast.txt --with-description --debug
    # → broadcast.txt + .stat.json (source="x", transcript_origin="macwhisper",
    #   chosen_track_kind="asr"). A status video WITH captions skips ASR
    #   (transcript_origin="embedded-captions"). Use --cookies-file for
    #   protected/age-gated media, or drop a Netscape cookies.txt at
    #   ~/.transcript-fetcher/x.com-cookies.txt for zero-flag auth
    #   (see the X cookie contract in §5 Safety Boundaries).
    
  • Skool lesson (cookies needed ONLY for private / paid communities — public ones work without):
    ./scripts/.venv/bin/python scripts/fetch.py \
        "https://www.skool.com/<community>/classroom/<id>?md=<lesson-id>" \
        --out lesson.txt --with-description \
        --cookies-file ~/.config/skool-cookies.txt
    
  • Batch:
    ./scripts/.venv/bin/python scripts/fetch.py --batch urls.txt --out-dir transcripts/
    
  • Inputs: A YouTube / Vimeo / Skool-lesson URL (or a file with one URL per line). Empty lines and # comments in the batch file are ignored.
  • Outputs:
    • <out>.txt — clean plain text, UTF-8.
    • <out>.txt.stat.json — sidecar with the chosen track, quality flag, plus optional title / uploader / upload_date / duration_sec / embed_source / embed_url metadata.
    • <out>.description.md — only when --with-description is passed; YAML frontmatter + Markdown body.
    • One JSON stat record per URL on stdout.
  • Failure semantics: Non-zero exit. With --json-errors, stderr carries a single JSON line {v, error, code, type, details?}. Exit codes: 2 usage error (incl. malformed Skool URL, cookies file path missing on disk), 3 no transcript producible (no caption track in the ladder AND — for X — ASR produced nothing / every available backend failed, or the media download timed out — that case is transient/retryable and its remediation names --concurrent-fragments / --media-timeout-sec), 4 partial batch failure, 5 source-auth error (HTTP 401/403 — private Skool community needs cookies, X protected/suspended/age-gated media, or supplied cookies expired), 6 source rate-limit (HTTP 429), 7 missing dependency (yt-dlp absent, ffmpeg required-but-absent, or no ASR backend available for caption-less media — details.remediation carries the hint), 1 unexpected. When details carries a remediation key, it is ALSO printed as a second, plain stderr line (remediation: <text>) even WITHOUT --json-errors — the remedy is operator-visible either way.
  • Idempotency: Re-running overwrites the output file and sidecar. yt-dlp itself caches nothing the skill depends on; behaviour is reproducible given network availability.
  • Dry-run support: Not currently exposed as a flag. Inspect the fallback ladder via _build_ladder if needed.

5. Safety Boundaries

  • Allowed scope: Reads from the network (yt-dlp HTTPS to YouTube/Vimeo; stdlib HTTPS to Skool). Writes ONLY to the user-specified --out / --out-dir path plus a .stat.json sidecar (and optionally a .description.md sidecar) next to it.
  • Default exclusions: Never downloads the video itself (only --skip-download + subtitle tracks / --write-info-json for metadata). Never writes to any path the user did not explicitly pass via --out / --out-dir.
  • Destructive actions: None. The script does not delete or modify any pre-existing files outside the chosen output paths. In batch mode, output path collisions are handled per --on-collision={error,skip,suffix} (default: error).
  • Optional artifacts: The JSON stat sidecar is mandatory in single-URL mode (it is the audit trail for which track was used). The .description.md sidecar is written only when --with-description is passed.
  • URL allowlist: Source dispatch is hostname-based against an explicit allowlist:
    • YouTube — youtu.be, youtube.com, m.youtube.com, music.youtube.com, youtube-nocookie.com (plus www. variants).
    • Vimeo — vimeo.com, www.vimeo.com, player.vimeo.com.
    • X / Twitter — x.com, www.x.com, mobile.x.com, twitter.com, www.twitter.com, mobile.twitter.com (status …/status/<id> and …/i/broadcasts/<id>).
    • Skool — skool.com, www.skool.com, app.skool.com; additionally URLs must match /<community>/classroom/<id>?md=<lesson-id>. Landing / /about / /calendar pages are rejected. URLs that merely contain a supported host as a substring elsewhere are rejected.
  • ASR backends (external, optional): For caption-less X media the skill shells out (argv arrays, never a shell string) to whichever local engine is present — MacWhisper mw, Whisper CLI, or whisper.cpp. None is a pip dependency; all are probed at runtime. ffmpeg (also external) is required to turn an X Broadcast's HLS stream into a valid audio file; the skill fails fast with exit 7 (clear remediation, before any large download) when ffmpeg is absent for an HLS source. bash scripts/install.sh reports which engines are available; scripts/install_components.py guides/installs them (incl. ffmpeg). If no ASR backend is available (and cloud is not opted in), the run also fails cleanly with exit 7, never a traceback. ASR portability: the fallback chain resolves in order mw → Whisper CLI → whisper.cpp → (opt-in) cloud (§2 Capabilities) — a caption-less Broadcast/Space genuinely REQUIRES ffmpeg and at least one of these; a box with neither (e.g. a bare Linux/CI runner) fails hard on exit 7. Remediate with scripts/install_components.py --install-whisper (in-venv Whisper CLI; ffmpeg itself needs the separate --system --run) or --asr-allow-cloud (+ an API key) to fall back to the cloud backend. Run scripts/fetch.py doctor (§4 Script Contract) before a long fetch to see which backends resolve, with zero downloads.
  • Cloud ASR egress (opt-in only): The OpenAI/compatible cloud backend is used only with --asr-allow-cloud (or TRANSCRIPT_FETCHER_ASR_ALLOW_CLOUD=1) AND an API key present. When used, the audio leaves the machine to the configured endpoint — disclosed here and in the stat notes. Local backends are always tried first; cloud is the last resort.
  • Silence removal before ASR (X): before transcribing, the X path runs ffmpeg silenceremove to trim leading silence and collapse long interior/trailing gaps — this cuts Whisper-family hallucinated filler (e.g. "Продолжение следует...") on silent lead-in/out. ON by default; --keep-silence (or TRANSCRIPT_FETCHER_SILENCE_REMOVAL=0) opts out; _THRESHOLD/_MIN_GAP_SEC/_KEEP_SEC tune it. Only true silence is removed (music/speech survive — a music-only intro can still trigger filler, see KNOWN_ISSUES TF-X-6). Never fatal: ffmpeg absent or a filter failure transparently falls back to the original audio. The stat notes record what was stripped (silence-removal: stripped ~Ns ...); the original media is kept for the ffprobe duration fill.
  • Secrets: The API key is read from OPENAI_API_KEY / TRANSCRIPT_FETCHER_OPENAI_API_KEY or a skill-local .env. A .env is loaded only at the CLI entry point and is refused if it is a symlink or not chmod 600 (group/world-readable). The key is sent only in an HTTP Authorization header — never on a command line, never logged. .env is git-ignored; only scripts/.env.example (placeholders) is committed.
  • Per-host cookies (~/.transcript-fetcher/): cookies for auth-walled media resolve (after an explicit --cookies-file) from a skill-local home folder — mirrors the html skill's ~/.html. An auth-map.json (--auth-map / TRANSCRIPT_FETCHER_AUTH_MAP / ~/.transcript-fetcher/auth-map.json) maps a host to its {cookies_file}, or the convention ~/.transcript-fetcher/<host>-cookies.txt is used — e.g. ~/.transcript-fetcher/x.com-cookies.txt for https://x.com/... URLs, ~/.transcript-fetcher/twitter.com-cookies.txt for https://twitter.com/... URLs. The convention lookup tries the EXACT URL hostname first, then — ONLY for the three well-known mirror-prefix labels www/mobile/m — the same file with that single label stripped (e.g. www.x.com and mobile.x.com both also resolve x.com-cookies.txt when the exact-host file is absent); distinct domains are never aliased to each other (x.com and twitter.com still need separate files — no generic parent-domain walk). A custom filename (e.g. x-cookies.txt) REQUIRES an auth-map.json entry; the convention path only matches the literal <host>-cookies.txt name (or its single-label-stripped mirror variant). Host match is label-boundary (a key x.com matches x.com/*.x.com, never evil-x.com); auth-map and convention files are hardened (symlink-reject + 0600). The resolved Netscape cookies.txt feeds yt-dlp's --cookies (and Skool's opener). --cookies-from-browser BROWSER loads cookies straight from a local browser via yt-dlp (opt-in — reads the browser's cookie store). On an X auth failure (SourceAuthError, exit 5) the message names the refresh path: the resolved --cookies-file when one was supplied, else the convention path to create — derived from the failing URL's own host (www./mobile. labels stripped), e.g. ~/.transcript-fetcher/x.com-cookies.txt for an x.com URL or ~/.transcript-fetcher/twitter.com-cookies.txt for a twitter.com URL; the convention lookup's mirror-prefix fallback above guarantees this hinted path is actually picked up on retry, for all 6 documented X hosts.
  • Temp-file hygiene: For X media, all intermediates (audio, VTT, info.json, .part, .m3u8, fragments) live under one tempdir removed in a finally block even on error — nothing is left behind.
  • Auth credentials: --cookies-file <path> accepts a Netscape cookies.txt and is ALWAYS OPTIONAL for every source. The file is read once at startup, never copied or re-emitted. For Skool, public communities (e.g. zero-one) serve lessons without auth; private / paid communities respond with HTTP 401/403 and the user then needs to supply cookies. YouTube/Vimeo optionally forward the file to yt-dlp's --cookies for age-gated or unlisted videos. The skill never blocks on missing cookies up-front — it tries the fetch and surfaces a SourceAuthError (exit 5) only if the source returns 401/403.

6. Validation Evidence

  • Local verification:
    cd skills/transcript-fetcher
    ./scripts/.venv/bin/python -m unittest discover -s scripts/tests
    
    All offline tests must pass without network. The end-to-end network test is gated behind TRANSCRIPT_FETCHER_E2E=1.
  • Skill-validator (structural):
    python3 .claude/skills/skill-creator/scripts/validate_skill.py skills/transcript-fetcher
    
  • Skill-validator (security):
    python3 .claude/skills/skill-validator/scripts/validate.py skills/transcript-fetcher
    
  • Expected evidence: Both validators exit 0; unittest reports OK.

7. Instructions

Step 1: Verify environment

If scripts/.venv/ does not exist, run bash scripts/install.sh first. The install script is idempotent — safe to re-run.

Step 2: Choose mode

InputModeCommand form
Single URLsinglefetch.py <URL> --out path.txt
List of URLs in a filebatchfetch.py --batch urls.txt --out-dir dir/

Step 3: Pick a fallback strategy

Default is --lang ru --prefer manual. This tries:

  1. manual:ru — user-uploaded Russian subtitles (highest quality).
  2. auto:ru-orig — YouTube auto-captions of the original Russian audio (good).
  3. auto:ru — YouTube auto-translation TO Russian (often noisy if speech was in another language).
  4. auto:en — English auto-captions as last resort (will set quality_flag = english_auto_translation).

For non-Russian content, pass --lang en (or another ISO code). For non-Russian languages the lang-orig step is skipped — it is a YouTube quirk that mainly matters for non-English speech.

Step 4: Run the CLI

Capture stdout (it carries the JSON stat). Read the stat to confirm which track was used:

./scripts/.venv/bin/python scripts/fetch.py \
    "https://youtu.be/NSVTpCfBMK8" \
    --out /tmp/talk.txt
# stdout: {"source":"youtube","url":"...","chosen_track_kind":"auto","chosen_track_lang":"ru-orig", ...}

Step 5: Inspect quality

Open the generated <out>.txt.stat.json. If quality_flag is set (currently only "english_auto_translation"), surface a warning to the user before passing the transcript to a downstream summarizer:

⚠️ TRANSCRIPT QUALITY: only English auto-translation was available for this URL. Idioms, proper names, and technical terms may be distorted. Consider asking the user for a manual transcription.

Step 6: Hand off

The clean .txt is now ready for summarizing-meetings or any other downstream consumer. Pass the path; do not paste the contents inline (transcripts are often large).

8. Workflows

- [ ] Verify scripts/.venv/ exists (run install.sh otherwise)
- [ ] Decide single vs batch mode
- [ ] Run fetch.py with the chosen language and preference
- [ ] Read the JSON stat sidecar
- [ ] Surface quality_flag warning if set
- [ ] Hand .txt path to downstream skill

9. Best Practices & Anti-Patterns

DO THISDO NOT DO THIS
Always read the stat sidecar after fetchingTrust the .txt without checking which track was used
Prefer ru-orig over ru for Russian contentPick the first track that returns text
Pass batch URLs through a fileLoop the CLI shell-side with arbitrary URLs
Surface quality_flag in any user-visible outputSilently downgrade to English auto-translation
Use --json-errors in CI/automation pipelinesParse free-form stderr

Rationalization Table

Agent ExcuseReality / Counter-Argument
"yt-dlp is on $PATH so I can just call it"The skill invokes python -m yt_dlp from the per-skill venv. The system yt-dlp may be a different version with different output.
"The >> markers are clutter"They are paragraph breaks for speaker turns. Downstream summarizers attribute statements by them.
"Rolling-caption dedup is overkill, just keep the longest cue"The dedup IS keeping the longest cue. Without it, the same sentence would appear 3-4 times.
"I'll add a Vimeo adapter inline in fetch.py"Add it as scripts/sources/vimeo.py. Each source is its own file.

10. Examples

See examples/:

  • example_input_url.txt — batch input format.
  • example_output_plain.txt — what a cleaned transcript looks like (excerpt).
  • example_output_stat.json — what the stat sidecar contains.

11. Resources

  • scripts/fetch.py — CLI entry point.
  • scripts/sources/youtube.py — YouTube adapter (yt-dlp orchestration + fallback ladder + description path).
  • scripts/sources/vimeo.py — minimal Vimeo adapter (yt-dlp).
  • scripts/sources/x.py — X.com / Twitter adapter (captions-first → ASR; the XTranscriptProvider).
  • scripts/sources/_ytdlp_media.py — shared yt-dlp plumbing (metadata probe, caption inspection, audio-minimal download, failure classifier) reused by X and any future yt-dlp source.
  • scripts/sources/_log.py — debug-only stage logger (stderr, gated on --debug).
  • scripts/sources/_auth.py~/.transcript-fetcher/ per-host cookie resolution (auth-map + convention, hardened; mirrors the html skill's ~/.html).
  • scripts/asr/ — pluggable ASR backend package: _base.py (the ASRBackend interface), macwhisper.py, whisper_cli.py, whisper_cpp.py, openai_api.py (opt-in cloud), __init__.py (priority registry + fallback chain).
  • scripts/_config.py — skill-local .env loader (secrets-safe) + typed config accessors (endpoints/models/tool paths).
  • scripts/.env.example — config/secret template (copy to .env, chmod 600).
  • scripts/install_components.py — detect / guide / install the optional ASR components.
  • scripts/sources/skool.py — Skool lesson adapter (cookies.txt + Next.js scrape + embed delegation).
  • scripts/sources/_vtt_to_text.py — pure-Python WebVTT cleaner.
  • scripts/sources/_captions.py — multi-format caption → text dispatch (SRT/TTML/DFXP build on the VTT cleaner; TTML XXE/billion-laughs guard).
  • scripts/sources/_stat.py — shared TranscriptStat + sidecar writer + error classes.
  • scripts/sources/_description.py.description.md writer (YAML frontmatter + Markdown body).
  • scripts/sources/_cookies.py — Netscape cookies.txt loader + authenticated opener.
  • scripts/sources/_prosemirror.py — ProseMirror/TipTap v2 JSON → Markdown for Skool lesson bodies.
  • scripts/install.sh — venv bootstrap.
  • scripts/requirements.txt — pinned deps (single source of truth for the yt-dlp version range).
  • scripts/tests/ — offline unit tests + opt-in E2E network test.
  • scripts/tests/_sanitize_fixture.py — utility for scrubbing PII from Skool HTML snapshots before they become fixtures.
  • references/youtube_caption_format.md — what >>, &gt;, rolling captions, and ru-orig actually mean.
  • references/fallback_policy.md — the language ladder and why it is in this order.
  • references/supported_sources.md — current and planned source slots.
  • references/skool_adapter.md — Skool auth flow, schema notes, embed delegation rules.
  • references/description_metadata.md.description.md format for YouTube and Skool.
  • docs/Manuals/transcript-fetcher_manual.md — user-facing manual with quick reference, troubleshooting, and composition recipes.

12. Composition

  • Composes well with summarizing-meetings: run transcript-fetcher first to get a clean .txt, then pass that file to summarizing-meetings for a structured Markdown summary. The two skills are intentionally separate — fetching is a deterministic file operation; summarizing is a prompt-first reasoning task.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.