Youtube
Extract transcripts from YouTube videos via the YouTube caption systemFrom its SKILL.md
npx -y skills add axoviq-ai/synthadoc --skill youtubeAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- runs commandsInstructs the agent to run 1 command, including `pip install youtube-transcript-api`.
What its file declares
Copied from the file, not written here
The file declares its own license as AGPL-3.0-or-later. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
2.5 KB, 529 tokens by cl100k_base, as published. Nobody here has run it
YouTube Skill
Extracts the transcript (captions) from a YouTube video using the YouTube caption system — no API key or audio download required. Optionally uses a vision-capable LLM to produce a structured summary (Overview, Topics, Key takeaways) before the raw transcript.
Setup
pip install youtube-transcript-api
Standalone usage
Transcript only (no LLM required):
import asyncio
from synthadoc.skills.youtube.scripts.main import YoutubeSkill
skill = YoutubeSkill() # no provider — returns raw timestamped transcript
async def main():
result = await skill.extract("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
print(result.text) # "[0:00] text [0:04] text ..."
print(result.metadata) # {"video_id": "...", "title": "...", "url": "..."}
asyncio.run(main())
With LLM summarization (pass any provider that implements complete()):
skill = YoutubeSkill(provider=my_provider)
result = await skill.extract(url)
# result.text contains:
# ## Executive Summary
# <Overview / Topics / Key takeaways>
#
# ## Transcript
# [0:00] ...
The provider must implement:
async def complete(messages, system=None, temperature=0.0, max_tokens=4096)
-> object with .text (str), .input_tokens (int), .output_tokens (int)
Message (used to build the messages list) is importable from
synthadoc.skills.base:
from synthadoc.skills.base import Message
When this skill is used
- Source starts with
https://www.youtube.com/,https://youtu.be/, orhttps://www.youtubekids.com/
To search YouTube by topic instead of ingesting a specific URL, use the web search skill — it filters Tavily results to YouTube domains automatically:
synthadoc ingest "youtube Moore's Law"
synthadoc ingest "youtube kids: Sesame Street"
synthadoc ingest "search for youtube: history of computing"
Limitations
- Only works for videos that have captions (auto-generated or manually added). If no captions are available the source is skipped with a warning.
- Private or deleted videos are skipped gracefully.
What ships with it: 3 files
6.5 KB alongside SKILL.md, 2 of them executable
scripts/
- __init__.pyruns0 B
- main.pyruns6.4 KB
- requirements.txt104 B