Social fetch
Fetch content from social-media URLs (HackerNews, Reddit, GitHub, X/Twitter, LinkedIn, YouTube, Bluesky, arXiv, Medium, Substack, RSS, generic articles) and run web/social searches (DuckDuckGo, Brave, SerpAPI, Tavily, X, HN, YouTube, Bluesky, arXiv) — output as clean markdown or structured JSON. Use whenever the user asks to "pull", "fetch", "download", "summarise", or "search the web/Twitter/HN/YouTube/Bluesky/arxiv" for content at a URL or query.From its SKILL.md
npx -y skills add jedi4ever/social-skills --skill social-fetchAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
3 things to look at
- reads credentialsReads from 2 credential sources: `./.env` and 1 more.
- 5 stars5 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 8 commands, including `scripts/social-fetch list` and 7 more.
SKILL.md
23.9 KB, ~5.5k tokens by cl100k_base, as published. Nobody here has run it
social-fetch skill
Wraps the social-fetch Go binary at scripts/social-fetch (relative to this skill).
Trust the CLI. It is the authority for every fetch and search supported by this skill. Always shell out to scripts/social-fetch — never reimplement fetching with WebFetch, curl, custom parsers, or hand-rolled API calls, even if the binary returns empty results or an error you find surprising. If a fetch comes back empty, surface that to the user and (if appropriate) re-run with --log - to see audit lines, but do not try to "fix it" by going around the CLI.
Before invoking any provider, run scripts/social-fetch list to see which platforms are usable in this environment. Each provider is tagged with one of three states:
[ok]— fully configured, fire away[!auth]— required env var not set; the row's suffix names which one (e.g.→ missing BRAVE_API_KEY). Do not use these providers — pick a different one in the same category. Suggesting the user configure the missing key is fine only if they explicitly ask how; don't proactively nag them about every unset key.[bridge]— needs the local browser bridge (LinkedIn / Medium / Substack / linkedin search & timeline). Only use afterscripts/social-fetch bridge statusreports connected.
The auto provider chains (-p auto) already skip unconfigured providers, so they're safe — but when an explicit provider name is needed (e.g. the user asks "search Twitter"), check list first to confirm [ok] status before invoking. For the MCP shape see social_fetch_list_providers — same data, structured as {name, status, missing} per category.
For platform-specific quirks, run scripts/social-fetch hints (no argument — dumps every platform's hints in one shot) before a search/fetch you haven't done recently. Captures things like "X recent search caps at 7 days strictly", "LinkedIn temp-bans accounts that scrape too fast", "Reddit anonymous search has worse relevance than tavily site:reddit.com". Pass a specific platform name to scope the output (e.g. hints x).
Subcommands
scripts/social-fetch fetch <url> [<url>...] [flags]
scripts/social-fetch search "<query>" [flags]
scripts/social-fetch timeline <user-or-url> [flags] recent activity for a user (X / LinkedIn)
scripts/social-fetch ask "<question>" [flags] grounded answer engine (perplexity / grok / openai / anthropic / gemini / tavily / serpapi)
scripts/social-fetch research "<question>" [flags] EXPERIMENTAL — multi-angle research (decompose → parallel fan-out → synthesize)
scripts/social-fetch bridge {start|stop|status|run}
scripts/social-fetch bookmarks {list|profiles} local browser bookmarks (chrome today; --platform NAME)
scripts/social-fetch hints [<platform>] per-platform quirks, rate limits, gotchas (x / linkedin / reddit / ...)
Run scripts/social-fetch --help for the full reference. Output defaults to markdown; pass -f json or -f jsonl for structured input to other tools.
Credentials (.env support)
Provider keys (X_API_KEY, X_API_SECRET, TAVILY_API_KEY, SERPAPI_KEY, BRAVE_API_KEY, PERPLEXITY_API_KEY, XAI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY/GOOGLE_API_KEY/GOOGLE_CSE_ID, YOUTUBE_API_KEY, BLUESKY_HANDLE/BLUESKY_APP_PASSWORD, GITHUB_TOKEN) and routing hints (HTML2MD_PROVIDER, HTML2MD_READER, YOUTUBE_TRANSCRIPT_PROVIDER, TAVILY_TOPIC) can be set in the shell or placed in a .env file. At startup the binary loads, in order:
./.env(current working directory)<binary_dir>/.env(sits next to the installed binary — typically~/.claude/skills/social-fetch/.env)
Already-exported shell vars always win over file entries.
Decision rules
- One URL → fetch it.
scripts/social-fetch fetch <url>auto-detects the source from the host (HN, Reddit, GitHub, X, RSS, or generic article). - Images / diagrams in the post → every fetched item carries a
Medialist (LinkedIn post photos, X/Twitter media, Medium/Substack figures, YouTube thumbnails, generic article images). The MCPsocial_fetch_fetchenvelope surfaces this as amedia[]array of{url, type, alt}entries. When the user asks "what's on the picture / diagram / screenshot in this post", the agent's vision-capable Read tool (Claude Code / Claude Desktop) can read eachmedia[].urldirectly — no extra fetch call. Emptyaltusually means the image is worth looking at; populatedaltis the author's caption. Same data is in the## Mediasection of the markdown content. - A list of URLs → batch. Pipe via stdin (
cat urls.txt | scripts/social-fetch fetch) or use-i FILE. Add-j 8for parallel fetches; output stays in input order. - Save to disk →
-o FILEfor one file,-o DIR/for one file per URL. - A user's recent posts → timeline.
scripts/social-fetch timeline <user-or-url> [-p x|linkedin] [--kind ...] [-n N]. Auto-detects the provider from URL; default for bare handles is X. See "Timeline subcommand" below. - A grounded question → ask.
scripts/social-fetch ask "<question>" -p perplexity|grok|openai|anthropic|gemini|tavily|serpapi. Returns synthesized answer + sources. Use this only when the user explicitly wants a synthesized answer; for raw documents usefetchorsearch. - A multi-angle research question → research (EXPERIMENTAL).
scripts/social-fetch research "<question>" --max-angles 5 --jobs 4. Decomposes into 3-8 angles, fans out parallel queries, synthesizes a final answer with citations. Use when you'd otherwise issue 4-8 manual queries. Costs roughly 2 LLM calls + N tool calls per question; useaskfor simple lookups instead. - A query → search. Pick the provider that matches the user's intent.
-p autowalksperplexity → tavily → brave → serpapi → duckduckgo; comma-lists like-p tavily,duckduckgodefine a custom fallback order. Each falls through on missing key / error / 0 results.- "search the web" / unspecified →
duckduckgo(no auth) - "search Brave" / privacy-focused web →
brave(needsBRAVE_API_KEY; native--last 7dvia freshness) - "high-quality web search for AI agents" →
tavily(needsTAVILY_API_KEY) - "Perplexity index without synthesis" →
perplexity(needsPERPLEXITY_API_KEY; same key asask -p perplexity, but cheaper since no LLM tokens) - "LinkedIn posts about <topic>" →
linkedin(requires the browser bridge + a logged-in LinkedIn session; up to 50 results per query via scroll-to-bottom + wheel-event lazy-load). Use sparingly. LinkedIn aggressively rate-limits accounts that scrape — running this back-to-back will get the user temp-banned. Prefertavily/perplexity/serpapifor general "who's writing about X" questions, and only reach for-p linkedinwhen LinkedIn-specific posts are explicitly the goal. - "search Bluesky" →
bluesky(no auth, native date filter) - "search arXiv" / academic papers →
arxiv(no auth, sorted newest-first) - "search HN" →
hackernews - "search Reddit" →
reddit(no auth, public search.json; rate-limited per IP) - "search Twitter/X" →
x(needsX_API_KEY+X_API_SECRET) - "search via Google" →
serpapi(needsSERPAPI_KEY) - "search YouTube" →
youtube(needsYOUTUBE_API_KEY; supports--last 7d/--afternatively, dates are strict)
- "search the web" / unspecified →
Flags worth remembering
| flag | when |
|---|---|
-f markdown|json|jsonl | format (default markdown) |
-o PATH | stdout / FILE / DIR/ |
-i FILE | URLs file (- = stdin; auto-detected when piped) |
-j N | parallel workers for batch fetch |
--no-comments | skip comment trees on HN/Reddit/X |
--max-comments N | cap comments per item |
--generic-extraction | force the catch-all article extractor (debug) |
--log - | print per-fetch audit lines to stderr |
Search-only:
| flag | when |
|---|---|
-p PROVIDER | pick search provider |
-n N | max results |
--after YYYY-MM-DD / --before YYYY-MM-DD / --last 7d | date filters |
--site DOMAIN / --exclude-site DOMAIN | domain filters (repeatable) |
Examples
# Pull a HN story with comments → markdown to stdout
scripts/social-fetch fetch https://news.ycombinator.com/item?id=43000000
# Pull a Medium article → structured JSON
scripts/social-fetch fetch https://medium.com/@alice/some-post -f json
# Batch from a file → one .md file per URL in ./out/
scripts/social-fetch fetch -i bookmarks.txt -o out/ -j 8
# Pipe a list → JSONL stream
cat urls.txt | scripts/social-fetch fetch -f jsonl > all.jsonl
# Search the web, last 7 days, restrict to two domains
scripts/social-fetch search "vercel ai sdk" --last 7d --site vercel.com --site ai-sdk.dev
# HN search — top stories about a topic
scripts/social-fetch search "rust async" -p hackernews -n 20
Timeline subcommand
scripts/social-fetch timeline <user-or-url> [flags]
-p PROVIDER x (default for bare handles) | linkedin
--kind KIND x: all (default), tweets, replies, retweets
linkedin: all (default), posts, comments, reactions
-n N max items (default 30)
--last DUR sugar for --after (e.g. 7d, 24h)
x has a hard 7-day cap
--after / --before yyyy-mm-dd or RFC3339
--expand (LinkedIn) re-fetch each item via the post fetcher (slow)
--no-reshares (LinkedIn) drop reposts from the timeline
User identifier accepts:
swyx(bare handle → x)@swyx(@implies x)https://x.com/swyx(auto-detected)https://www.linkedin.com/in/patrickdebois/(auto-detected)patrickdebois+-p linkedin
LinkedIn timelines drive the bridge through scroll/get_html cycles. Returns ~5–50 items depending on how active the user is. Bridge required for LinkedIn — check bridge status first. X timelines wrap recent-search; the 7-day cap applies and the binary pre-flights it with a clear error.
# Last 7d on X
scripts/social-fetch timeline swyx --last 7d
# LinkedIn posts only (no reshares), markdown
scripts/social-fetch timeline patrickdebois -p linkedin --kind posts --no-reshares
# LinkedIn full deep-fetch (each item gets its body + comments)
scripts/social-fetch timeline matthewskelton -p linkedin --expand -n 10
Ask subcommand
scripts/social-fetch ask "<question>" [flags]
-p PROVIDER perplexity (default), grok, openai, anthropic, google, tavily, serpapi
special values:
auto try the built-in chain in order
(perplexity → grok → openai → anthropic →
google → tavily → serpapi)
name1,name2,… comma-list to try in order
-m MODEL override the provider's default (empty = provider picks where supported)
--last WINDOW day | week | month | year (provider-dependent)
--max-tokens N cap response length
--instructions system-prompt-style preamble (alias: --system)
honored by perplexity / grok / openai / anthropic / google;
ignored by tavily / serpapi (no system-prompt support)
Returns a synthesized answer plus a numbered Sources list. Auth needed per provider — see Credentials above.
When -p auto or a comma-list is given, each provider in turn falls through on (a) missing API key, (b) upstream error, or (c) empty answer — the next provider gets a try, and the first non-empty response wins. The audit log records which provider answered.
scripts/social-fetch ask "what changed in the openai-microsoft revenue share clause" -p grok
scripts/social-fetch ask "best agent harness papers in the last month" -p perplexity --last month
scripts/social-fetch ask "what's the weather in NYC" -p auto # try the default chain
scripts/social-fetch ask "what's the weather in NYC" -p perplexity,anthropic,duckduckgo # custom chain
Listing supported sources/providers
scripts/social-fetch list
Transports — bridge / headless / http / jina
The fetcher walks a chain of transports per platform; each fetched item carries Extra.via naming which one produced the body. Defaults live in the platform Go code; override per call via SOCIAL_FETCH_CHAIN_<NAME>. Run scripts/social-fetch hints <platform> for the per-platform recipe.
| transport | what it does | when needed |
|---|---|---|
bridge | drives your real, logged-in Chrome via the extension | auth-walled content (LinkedIn comments, Medium / Substack member-only posts) |
headless | local stealth Chromium via chromedp; anonymous, JS-rendering | JS-rendered SPAs, soft anti-bot — also the article fetcher's preferred path |
http | plain HTTP GET | static pages where JS isn't needed |
jina | remote r.jina.ai service | last-resort catch-all when local methods fail |
Defaults (mirrored from code — social-fetch hints <platform> is canonical):
| platform | default chain |
|---|---|
| article | headless,http,bridge,jina |
headless,bridge,jina (bridge for comments) | |
| medium / substack | bridge,http,headless,jina (bridge for paywall) |
api,syndication,jina |
LinkedIn requires the bridge
Setup once: load extensions/chrome/ (at repo root) as an unpacked Chrome extension.
Bridge lifecycle:
scripts/social-fetch bridge start # daemonize, write PID file
scripts/social-fetch bridge status # connected / not connected / not running
scripts/social-fetch bridge stop # graceful SIGTERM
scripts/social-fetch bridge run # foreground (good for `nohup` or terminals)
Always check status before fetching authenticated URLs:
$ scripts/social-fetch bridge status
connected # → fetch will work
not connected # → bridge up but extension hasn't attached (open the browser)
bridge not running on :5555 # → run `bridge start` first
Exit codes are 0 connected / 1 not connected / 2 bridge not running, so agents can branch on them.
Then fetch:
scripts/social-fetch fetch https://www.linkedin.com/posts/foo-activity-700…
The bridge tells the extension to navigate the URL in your real browser, scrapes the rendered DOM, and returns clean markdown.
URLs the LinkedIn fetcher claims: linkedin.com/posts/…, linkedin.com/feed/update/urn:li:activity:…, linkedin.com/in/<user>, linkedin.com/pulse/….
Errors you may see:
bridge unreachable→ start it (bridge start).no extension connected→ open your browser; the extension reconnects every ~6s.
Local browser bookmarks (scripts/social-fetch bookmarks)
Reads Chrome's local Bookmarks JSON and lists matching entries. Date-range filterable, multi-profile aware.
scripts/social-fetch bookmarks list # newest 100, default profile
scripts/social-fetch bookmarks list --since 2026-04-01 # bookmarked since April
scripts/social-fetch bookmarks list --folder-contains AI -n 20 # fuzzy folder match
scripts/social-fetch bookmarks list --folder "Bookmarks bar/AI" # exact subtree (AI/, AI/papers/, …)
scripts/social-fetch bookmarks list --all-profiles -f json # every profile, JSON
scripts/social-fetch bookmarks profiles # which profiles exist
--platform chrome is the default. Future platforms (Twitter bookmarks, Reddit saved posts — server-side, account-scoped) will plug in as additional values.
Scope every call to one folder via env var: set
SOCIAL_FETCH_BOOKMARKS_ROOT_FOLDER="Bookmarks bar/AI" once and the
agent's bookmarks list calls (CLI + MCP) only see bookmarks under
that folder + every nested subfolder. Override per-call with --folder.
Ledger daemon (sandboxed / remote MCP)
social-ledger daemon start daemonises the SQLite ledger behind an HTTP API on port 5557. When it's running, every caller — CLI, social-fetch's auto-ingest, MCP read tools — routes through HTTP instead of opening the SQLite file directly.
scripts/social-ledger daemon start
scripts/social-ledger daemon status
scripts/social-ledger daemon stop
In daemon mode, social_fetch_fetch returns content_url (HTTP pointer) instead of content_file (local path). Agents that don't have filesystem access to the daemon's host can still read fetched bodies. For local single-machine usage, leave it off — direct file access is faster (~10ms saved per call).
Multi-project ledgers: every ledger lives under <base>/projects/<NAME>/. The default project is social_fetch; set SOCIAL_LEDGER_PROJECT=research-x to switch to a separate ledger for that context. Pre-projects bare ledgers migrate automatically on first run (no operator action needed). One daemon serves one project; run multiple daemons on different ports for parallel projects.
Headless browser pool (anonymous JS-rendered fetches)
Separate from the bridge. social-browser daemon start --provider local daemonises a pool of warm headless Chromium browsers. Article / LinkedIn / Medium / Substack chains include headless and route through the daemon; typical 1–3s when warm. Anonymous-only — no session reuse, never touches the user's real Chrome profile.
scripts/social-browser daemon start --provider local --pool-size 2 --recycle-after 50
curl -s http://127.0.0.1:5560/status # snapshot
scripts/social-browser daemon stop
There is no in-process fallback — when no daemon is running, social-fetch screenshot (and any JS-rendered platform chain) returns a clean error pointing at the start command above. No silent slow-path.
For remote pools (chromedp running in Daytona sandboxes, fronted as a round-robin proxy):
scripts/social-browser provider daytona up -n 3
scripts/social-browser daemon start --provider daytona
When to start it: before any batch fetch with -j > 1, before research loops, when fetching JS-rendered articles where http returns a thin shell, before social-fetch screenshot.
YouTube
scripts/social-fetch fetch <youtube-url> claims youtube.com/watch?v=…, youtu.be/…, youtube.com/shorts/…, youtube.com/live/…, youtube.com/embed/…, and music.youtube.com/….
- Metadata: pure scraping via
kkdai/youtube/v2— no auth needed. - Transcript: configurable provider chain (see below). Appended to
Contentunder a## Transcriptheading; structured timed segments live inExtra.transcript. - Comments: optional, gated on
YOUTUBE_API_KEY(free Google Cloud key, 10,000 units/day). Without it, comments are skipped silently.
Transcript provider switching
Set YOUTUBE_TRANSCRIPT_PROVIDER to control which transcript backend is used:
| value | behavior |
|---|---|
auto (default) | yt-dlp if installed → InnerTube (no auth) → kkdai. First success wins. |
ytdlp | shells out to yt-dlp (most reliable; install with brew install yt-dlp or pip install yt-dlp) |
innertube | pure-Go scrape via youtubei/v1/get_transcript — fragile (YouTube can break it) but no extra runtime dep |
kkdai | the kkdai library's caption-track endpoint; YouTube has been gating this with HTTP 400s in 2026 |
Set YOUTUBE_API_KEY in your shell or .env for comments. Some videos have transcripts disabled by the channel — the fetcher logs that case and returns metadata + comments only.
Tavily date filter caveat
Tavily's general topic (the default — high relevance) doesn't populate published_date for most results, so --last 7d / --after enforce date strictly only on results we can date. Set TAVILY_TOPIC=news (in env or .env) when you want a guaranteed window — that switches Tavily's index to news-only, which has dates upstream + much narrower recall (often unhelpful for personal-name or evergreen-topic queries).
X / Twitter reply behavior
When X_API_KEY + X_API_SECRET are set, fetching a tweet also pulls its replies as a nested tree (one batched tweets/search/recent call per 100 replies — no per-reply round-trips). Caveats:
- Search is limited to the last 7 days by X's API tier — older tweets return 0 replies. The audit log (
--log -) makes this explicit. - Without creds, the syndication fallback is used and returns 0 replies (no API support).
--no-commentsand--max-comments Napply.
Ledger (history of every fetch)
A companion binary scripts/social-ledger ships alongside
scripts/social-fetch. When both are present (always true for this
skill), every successful fetch / timeline / research call is
auto-recorded — content into a SQLite + FTS5 store, plus a
markdown mirror tree on disk. No env-var setup needed; auto-detect
flips it on.
Useful queries against the ledger from inside the skill:
scripts/social-ledger article list # newest first
scripts/social-ledger article list --source hackernews # filter by source
scripts/social-ledger article search "harness engineering" # full-text search
scripts/social-ledger article get https://example.com/foo # one item back
scripts/social-ledger article stats # counts + sizes
scripts/social-ledger article forget https://... # drop one entry
When the user asks a question that resembles "have we seen X before?" or "what did we learn about Y last week?" — the ledger is where to look first, before re-fetching.
To explicitly disable: set SOCIAL_LEDGER=0 in the env. To
override the storage location: SOCIAL_LEDGER_DIR=...
(default ~/.local/share/social-ledger).
When NOT to use this skill
- The user wants to post content (this skill only reads).
- The URL is behind a paywall/login — output will be the gated stub. Tell the user.
- The URL needs a logged-in browser session (LinkedIn, X home feed, etc.) — not supported.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most social media skills give in ~5.5k tokens
Counted across 400 of the 422 authors here whose files we hold, read 2026-09-06
- Build content around three to five pillarsin 40 of 400, across 16 files
- Read product marketing context before asking questionsin 37 of 400, across 15 files
- Gather goals, audience, brand voice, and resourcesin 35 of 400, across 13 files
- Maintain one to two weeks of scheduled contentin 27 of 400, across 8 files
- Batch content creation in weekly sessionsin 23 of 400, across 7 files
- Review top and bottom posts weeklyin 22 of 400, across 7 files
- Run the quality gate before deliveringin 20 of 400, across 11 files
- Add subtitles to all social videoin 20 of 400, across 7 files
- Write standalone captions that work without contextin 20 of 400, across 7 files
- Prefer specificity over adjectivesin 19 of 400, across 10 files
- Respond to all comments on your postsin 18 of 400, across 6 files
- Carry one actual claim per postin 18 of 400, across 9 files
Said here and by no other author read
- Use the social-fetch CLI for every fetch and search
- Run list before invoking any provider
- Run hints before an unfamiliar fetch or search
- Check bridge status before fetching authenticated URLs
- Surface empty fetch results to the user
- Pick the search provider matching the user's intent
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.