agentsclimarketplace

Yt structure

Skill Pesty-Marketing/agent-skills/yt-structure

Use when turning a YouTube video, talk, podcast, or web article into a clean, citable Markdown document for an LLM knowledge library or RAG/OKF bundle — triggers like "transcribe and structure this video", "structure this transcript", "add this talk/article to the library". Covers pulling captions, faithful structuring, and editor's-note / [?] flagging.From its SKILL.md

Install
npx -y skills add Pesty-Marketing/agent-skills --skill yt-structure

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

5.2 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it

yt-structure

Overview

Turn a spoken or web source into a faithful, scannable, citable Markdown file for LLM ingestion.

Core principle: structure and clean phrasing only — never paraphrase substance, never invent content, and flag what you can't verify instead of guessing.

When to use

  • A YouTube video, conference talk, podcast, or lecture → a cleaned .md.
  • A blog post / web article you want as faithful text (not a summary).
  • Adding a source to your knowledge library (a target library folder, an OKF/RAG bundle — at Pesty this is usually the internal Agents library).

When NOT to use:

  • A messy meeting transcript with crosstalk (Gemini/Fathom) → that's a separate internal cleanup process (de-fragment turns), not covered by this skill.
  • You just want a quick summary — this produces a faithful structured document, not a digest.

Step 1 — Ingest the raw source

Video / talk / podcast:

Use the bundled script at scripts/yt-transcript (relative to this skill folder):

scripts/yt-transcript <youtube-url> [output-dir]

Writes a clean <slug>.txt (strips YouTube's rolling-window caption duplication) and prints the path. Run it with no args for usage. Prerequisite: yt-dlp (brew install yt-dlp or pipx install yt-dlp). Do NOT use Tactiq, web tools, or youtube-transcript-api — they waste tokens and break sentences.

If the script isn't available in this environment, fall back to the raw command it wraps:

yt-dlp --skip-download --write-auto-sub --sub-lang 'en.*,en' --sub-format vtt -o '%(id)s.%(ext)s' <url>

then strip the timestamp/tag markup and deduplicate repeated caption lines yourself.

Web article:

curl -sL "<url>" | pandoc -f html -t markdown_strict

Use this for blog posts — agent web-fetch tools summarize prose and lose faithful text. (Official docs pages usually come through a web-fetch tool fine.) Prerequisite: pandoc (install via brew/apt) — or any equivalent HTML→Markdown converter works.

Step 2 — Structure into a sibling .md

Write <slug>.md next to the raw source, same slug. Follow these conventions exactly:

  • Header block: # Title, then **Speaker:** / **Author:** line, then **Source:** <url>.
  • ## headers at topic transitions; ### for sub-beats.
  • Strip filler (um, uh, "kind of", "you know", redundant so/well/right, doubled phrases). Phrasing cleanup only.
  • > [Visual: …] blockquotes for on-screen references (clips, slides, diffs, B-roll, demos).
  • Verbatim quotes / cited passages → blockquotes; bold the takeaways.
  • End with ## Recap.
  • ### Editor's note at the end: list proper-noun corrections you made + any unresolved [?] flags.
  • Multi-speaker: preserve who-said-what; never invert authorship (don't credit a listener/agreer with the speaker's point).
  • Flag, don't guess: garbled names/tools/companies get a [?]. Verify the big factual flags (names, tools, claims) with a web search before de-flagging.

Example output (illustrating the conventions above):

# Why Most SaaS Onboarding Fails
**Speaker:** Jane Rivera
**Source:** https://youtube.com/watch?v=abc123

## The core problem

### Users never reach the "aha moment"
**Most churn happens before value is ever felt.** Rivera argues teams over-invest
in visual polish and under-invest in the first five minutes of use.

> [Visual: funnel chart showing 80% drop-off before day two]

> "We rebuilt onboarding around one metric: time to first real result."

She attributes the framework to Kathy Sierra [?] — spelling unconfirmed.

## Recap
Ship a narrow, guided first-run path; measure time-to-value, not signups.

### Editor's note
Corrected "Rivera" (garbled as "Rivero" in captions). Unresolved: `[?]` on
"Kathy Sierra" — not yet verified via web search.

Step 3 — Add to the library (if building one)

Append the new file to the library's INDEX.md / index.md. If the library is an LLM-wiki / OKF bundle, also update the concept pages the source touches (don't just add a file) — follow that library's CLAUDE.md maintenance loop.

Common mistakes

MistakeFix
Tactiq / youtube-transcript-api for captionsUse scripts/yt-transcript (dedups, LLM-ready)
A web-fetch tool for a blog articleIt summarizes — use curl … | pandoc for faithful text
Paraphrasing or compressing substanceClean phrasing only; keep the content
Guessing a garbled proper noun[?] flag + a web search to confirm
Inverting authorship in multi-speaker talksPreserve attribution; agreement ≠ authorship
Skipping the Recap / Editor's noteBoth are required closers
</content>

Maintained by Pesty Marketing · Browse the full skill catalog.

What ships with it: 1 file

1.4 KB alongside SKILL.md, 1 of them executable

scripts/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.