agentsclimarketplace

Podcast transcript fetcher

Skill Varnan-Tech/opendirectory/skills/podcast-transcript-fetcher

AI Agent Skills built for Founders who hate Marketing

Install
npx -y skills add Varnan-Tech/opendirectory --skill podcast-transcript-fetcher

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Use when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and most reliable. Also use when asked to "get transcript", "find episode", "summarize podcast", or "search podcast content". Do not use for general web scraping or non-podcast audio transcription.

SKILL.md

7.5 KB, as published. Nobody here has run it

Podcast Transcript Fetcher

Fetch transcripts from 5 supported podcasts. Tier 2 (RSS+Groq Whisper) is the recommended approach -- fast, free, and the most reliable across all podcasts. Tier 1 free sources are best-effort (limited availability). Tier 3 Taddy API is the premium/commercial option.

Quick Reference

# Get latest episode transcript (auto-detects best method)
python scripts/get_transcript.py "Lenny's Podcast" --latest

# Search by episode title or number
python scripts/get_transcript.py 20vc --episode "Marc Andreessen"
python scripts/get_transcript.py dwarkesh --episode 15

# Force specific method
python scripts/get_transcript.py "cheeky pint" --latest --method whisper
python scripts/get_transcript.py a16z --latest --method taddy

# Save to file
python scripts/get_transcript.py lennys --latest --output transcript.md

# List all supported podcasts
python scripts/get_transcript.py --list-podcasts

Supported Podcasts

PodcastTier 1 (best-effort)Tier 2 RSS+Whisper [RECOMMENDED]Tier 3 Taddy (premium)
Lenny's PodcastGitHub archive (269 transcripts)✅ Substack RSS✅ Covered
Dwarkesh PodcastWebsite scrape + Substack PDF✅ Substack RSS✅ Covered
Cheeky Pint(none)✅ Transistor.fm RSS✅ Covered
20VCSubstack PDF✅ Libsyn RSS✅ Covered
A16z PodcastWebsite scrape✅ Simplecast RSS✅ Covered

Implementation

1. Install Dependencies

# Core (always required)
pip install requests

# Cloud transcription (recommended — fast, free tier)
pip install groq
export GROQ_API_KEY="your-key"  # Get at https://console.groq.com

# Local transcription (free, needs ~5GB RAM)
pip install faster-whisper

# Audio compression (for Groq's 25 MB limit — Windows: winget/scoop)
#   winget install ffmpeg  or  scoop install ffmpeg

# Taddy API (commercial, optional)
export TADDY_API_KEY="your-key"  # Get at https://taddy.org

2. Get a Transcript

The script auto-selects the best method. Tier 2 is the default recommendation:

Tier 1 → Tier 2 (RECOMMENDED) → Tier 3
(best-effort)  (Whisper)  (Taddy API premium)

Tier 1: Free direct sources (best-effort, limited availability)

  • Lenny's: Clones ChatPRD/lennys-podcast-transcripts and searches by title
  • Dwarkesh: Substack PDF scrape
  • 20VC: Substack PDF scrape
  • A16z: Website scrape
  • Cheeky Pint: No Tier 1 sources available
  • Note: Tier 1 sources are best-effort and limited. Tier 2 (RSS+Whisper) is the recommended approach.

Tier 2: RSS + Whisper transcription [RECOMMENDED]

  • Downloads MP3 from podcast RSS feed
  • Compresses if >25 MB (ffmpeg)
  • Transcribes via Groq Whisper API (free tier, ~10s per hour of audio)
  • Fast, free, and works for every podcast in the registry
  • Default recommendation for all use cases

Tier 3: Taddy API (commercial/premium)

  • Requires TADDY_API_KEY ($75/mo+)
  • Use for large-scale or production transcript needs
  • Covers all 5 podcasts with auto-transcription

3. Analyze with AI

Once you have a transcript, pipe it to the agent for analysis:

I have this transcript from [podcast]. Can you:
1. Summarize the key arguments
2. Extract 3 actionable insights
3. Identify any controversial claims
4. Compare with [other podcast] on the same topic

Supported Workflows

Single Episode

ScenarioCommand
Latest episodeget_transcript.py "Lenny's Podcast" --latest
Specific episode by titleget_transcript.py 20vc --episode "Sam Altman"
Episode by numberget_transcript.py dwarkesh --episode 42
Force Whisper transcription (Tier 2, recommended)get_transcript.py a16z --latest --method whisper
Force Taddy API (premium)get_transcript.py lennys --latest --method taddy
Save to Markdownget_transcript.py cheeky-pint --latest --output episode.md
JSON outputget_transcript.py dwarkesh --latest --json

Cross-Podcast Search & Batch

ScenarioCommand
Search all podcasts by keywordget_transcript.py --search "Marc Andreessen"
Search by guest nameget_transcript.py --guest "Sam Altman"
Search within one podcastget_transcript.py "Lenny's Podcast" --search "vibe coding"
Batch-transcribe last N episodesget_transcript.py "Dwarkesh Podcast" --last 5
Search + transcribe top matchesget_transcript.py --search "AI safety" --transcribe
Pipeline with custom countget_transcript.py --search "scaling laws" --transcribe --transcribe-count 5
Filtered search pipelineget_transcript.py "A16z Podcast" --search "crypto" --transcribe

Output Structure

Batch transcription saves to output/ with per-podcast subdirectories:

output/dwarkesh-podcast/Dwarkesh Podcast_2024-01-15_agi-is-still-30-years-away.md
output/20vc/20 Minutes VC (20VC)_2024-03-10_funding-round-analysis.md

Each file includes a YAML frontmatter header:

---
podcast: Dwarkesh Podcast
episode: AGI is still 30 years away
date: 2024-01-15
url: https://...
source: whisper
---

Podcast Registry

The registry at scripts/podcasts.json maps each podcast to its RSS feeds, transcript sources, and API endpoints. To add new podcasts:

{
  "id": "new-podcast",
  "name": "New Podcast",
  "rss": "https://example.com/feed.xml",
  "transcript_sources": {
    "primary": {"type": "website_scrape", "url": "https://example.com"}
  }
}

Troubleshooting

ProblemSolution
"No transcript found"Tier 2 (RSS+Whisper) is the recommended approach. If auto mode fails, try --method whisper to force it.
RSS fetch failsRSS feeds may change; check scripts/podcasts.json for current URLs
Audio download slowLarge MP3s can take minutes on slow connections
Groq rate limitedWait or switch to local faster-whisper
Taddy not returning transcriptsSome episodes lack transcripts; try --method whisper
Podcast not in registryAdd it to scripts/podcasts.json
Unicode error on WindowsFixed: script auto-reconfigures stdout to UTF-8; saved files use UTF-8 encoding
Audio > 25 MB for GroqInstall ffmpeg: winget install ffmpeg (Windows) or brew install ffmpeg (macOS)

RSS Feed Status (as of 2026-06)

PodcastOld Feed (broken)Current Feed
Cheeky Pintfeeds.transistor.fm/the-cheeky-pint (404)feeds.transistor.fm/cheeky-pint-with-john-collison
20VCfeeds.simplecast.com/3GxrMqOd (404)feeds.libsyn.com/61840/rss
A16zfeeds.simplecast.com/0cJfpoz2 (404)feeds.simplecast.com/JGE3yC0V

Common Mistakes

  • Forgetting API keys: Set GROQ_API_KEY in your env or .env file
  • Relying on Tier 1 free sources: Tier 1 is best-effort and limited. Always fall back to Tier 2 (RSS+Whisper) which is the recommended method.
  • Not cloning the Lenny's repo first: The GitHub archive must be cloned locally for Tier 1 to work
  • Using --method taddy without TADDY_API_KEY: Falls through silently; set the key or use auto mode

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.