Clipsmith x
Download X (Twitter) post assets — text, images, video — to a local folder using browser automation with manual-login session reuse. Use when downloading an X post, saving post images, exporting post content, or archiving a tweet. Trigger phrases: "download x post", "save tweet", "x post images", "x video download", "twitter post assets".From its SKILL.md
npx -y skills add OctopusGarage/clipsmith --skill clipsmith-xAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
9.0 KB, ~2.0k tokens by cl100k_base, as published. Nobody here has run it
clipsmith-x
MANDATORY — load
references/plan.mdbefore any browser or extraction action begins.
⚠️ NEVER WRITE YOUR OWN SCRIPT
The download logic is fully implemented. Always invoke the existing script — do NOT write a new one.
cd /Users/kingsonwu/programming/OctopusGarage/clipsmith/skills/clipsmith-x
npx tsx scripts/run.ts \
--post-url "<url>" \
--output-dir "$HOME/Downloads/x"
The sections below (image extraction, anti-detection, MHTML generation) are implementation documentation for the script itself, not instructions for you to re-implement. If the script doesn't exist or can't run, report the error — never substitute with hand-written Playwright code.
Clipsmith Bundle Normalization
The copied downloader produces a raw post folder with post.md, images,
optional video, and optional MHTML. Before finalizing a Clipsmith capture job,
convert that raw folder into a bundle with the shared normalizer:
cd /Users/kingsonwu/programming/OctopusGarage/clipsmith
uv run clipsmith normalize raw x "<raw_dir>" "<bundle_dir>" \
--source-url "<original_url>" \
--canonical-url "<canonical_url>" \
--title "<title>" \
--author "<author>" \
--captured-at "<iso8601_time>" \
--json
uv run clipsmith validate-bundle "<bundle_dir>" --json
The normalizer keeps post.md, creates or copies summary.md, preserves
ocr.md/ocr.txt as kind: "ocr-text" if present, and writes capture.json.
It intentionally does not copy downloaded media or MHTML into the final bundle
because the bundle validator does not allow arbitrary raw assets.
Do not call uv run clipsmith capture finalize until capture.json exists and
validation succeeds.
Quality Evaluation
Use the committed eval profile and fixture before changing prompt, extraction, media, MHTML, post-type detection, t.co expansion, or normalization behavior:
cd /Users/kingsonwu/programming/OctopusGarage/clipsmith/skills/clipsmith-x
node scripts/eval.mjs \
--fixture x-kingson-skill-runtime-text \
--profile x-kingson-skill-runtime-text
When live X access is available, also validate the known media and article profiles:
node scripts/eval.mjs \
--post_dir "/path/to/x/media-output-folder" \
--profile x-droidbuilds-single-image
node scripts/eval.mjs \
--post_dir "/path/to/x/article-output-folder" \
--profile x-geekcatx-chatgpt-codex-article
The committed fixture is a user-owned text-only post. External media and article
profiles are live eval profiles only; do not commit third-party full text,
images, video, or MHTML unless permission is explicit. Use
prompts/evaluate-capture.md for agent AI eval and compare against the fixture
baseline evals/ai-evals/x-kingson-skill-runtime-text.md when working in the
source repo. Packaged skill installs may omit eval fixtures and baselines.
URL Expansion (t.co Shortlinks)
Tweet links render as t.co shortlinks (e.g., https://t.co/yr4YXZ6SgU) — both in href attributes and textContent. post.md must output the full resolved URL (e.g., https://github.com/anthropics/skills/tree/main/skills/pptx).
Because Twitter's Content Security Policy (CSP) blocks cross-origin fetch() inside page.evaluate(), URL expansion must be done in Node.js:
- Collect all
<a>elementhrefattributes insidepage.evaluate() - Call
fetch(href, { redirect: "follow" })onhttps://t.co/*links from Node.js to get the final URL - Replace each anchor's
textContentwith the resolved URL, then extract full text
This step is integrated into extractTweetSnapshot() and must not be skipped.
Required Constraints
- Use browser automation only.
- Do not use X private APIs.
- Reuse manual-login session via unified Chrome CDP startup:
open -na "Google Chrome" --args --remote-debugging-port=9222 --user-data-dir="$HOME/.chrome-labali" --no-proxy-server. - Prefer semantic extraction from visible page state and loaded resources.
- Download only target post assets: images plus optional post video.
- Generate
post.mdfor extracted text metadata. - Do not generate
manifest.json. - Preserve all query parameters for page navigation.
Anti-Detection Principles
X applies behavioral analysis to detect automation. Violations of these principles may result in account restrictions.
Core test — apply before every browser action:
"Would a real user do this, from this state, at this moment?" If no → skip it or slow it down. Re-navigating an already-open post, batch-extracting DOM nodes, issuing fresh HTTP requests for images the browser just loaded, using fixed delays — all fail this test.
Navigation:
- If the tab is already on the target post URL, skip
page.goto()entirely. - All fixed
waitForTimeoutvalues must be randomized (e.g.,base + Math.random() * range). - After navigating to a post, scroll down briefly to simulate reading, then scroll back up before interacting with images.
Image acquisition:
- Never issue new HTTP requests for images — the browser has already downloaded them.
- Register
page.on("response")BEFORE callingpage.goto()— images load during navigation. - Primary image source: DOM extraction from
[data-testid="tweetPhoto"]— handles both<img>elements (regular tweets) and CSSbackground-imagedivs (X Notes). - Fetch each URL via
fetch(url, {cache: 'force-cache'})inpage.evaluate(). - After download: deduplicate by SHA-256 content hash; remove files < 40% of median size.
MHTML Generation (CDP Page.captureSnapshot):
- Calls Chrome's built-in
Page.captureSnapshotcommand via Playwright CDPSession to generate standard MHTML - Before archiving, automatically removes UI elements unrelated to post content (login buttons, nav bars, recommended content, etc.)
cleanupPageForArchiveremoves: LoginForm, signup links, banner nav, sidebarColumn, app-bar, etc.- Text-only tweets (no images/video) skip MHTML and only generate
post.md - X Notes and posts with images/video generate
article.mhtml+post.md
Video:
- Before downloading, simulate user engagement: bring the tab to front, click the video element, wait 3–5 seconds (randomized) for buffering.
- Video download uses
page.request.get().
General:
- Always operate within the user's authenticated Chrome session (CDP reuse) — never launch a headless or separate browser.
- Never manipulate the DOM beyond what a user's own browser JS would do.
NEVER
- Never launch a new Chrome if CDP is already responding on port 9222.
- Never hijack a non-X browser tab — reuse X tab or open a new X tab.
- Never strip query params from post URL before navigating.
- Never use fixed (non-randomized) delays.
- Never retry after CAPTCHA or rate-limit signal.
- Never report success without verifying output files exist.
Success Criteria
A run is successful only when all conditions hold:
- An output folder is created:
<output_dir>/<YYYYMMDD>-<tweet_id>/ post.mdis generated in the folder with author, text, timestamp, URL.- Post image files are saved as
image_01.jpg,image_02.jpg, etc. - Post video file is saved as
video.mp4when the post contains video. article.mhtmlis generated for X Notes and posts with images/video; text-only tweets omit it.- Files are deduplicated by content hash.
- URL output and logs use canonical
x.comform.
Operational Mode
- Default: guided browser flow + semantic extraction + authenticated media download.
- Startup:
- Check if CDP responds:
curl -s http://localhost:9222/json/version - If CDP NOT responding → auto-launch Chrome immediately:
open -na "Google Chrome" --args --remote-debugging-port=9222 --user-data-dir="$HOME/.chrome-labali" --no-proxy-server; wait 3s; verify CDP responds - If Chrome with remote debugging is already running, reuse it
- Find existing x.com tab → reuse it; if none → open new tab
- Check login state; if wall detected → guide user to complete login manually; continue in same session
- Check if CDP responds:
- Input:
- After startup, if
post_urlis missing → prompt user interactively - If
output_diris missing → prompt with default~/Downloads/x
- After startup, if
Resources
| When | Must load | Do NOT load |
|---|---|---|
| Always — at skill invocation start | references/plan.md | references/architecture.md |
| Extraction returns wrong count or fails | references/architecture.md | — |
| Video download unclear | references/architecture.md | — |
See Resources table above for conditional loads.
What ships with it: 23 files
82.9 KB alongside SKILL.md, 5 of them executable
agents/
- openai.yaml1014 B
evals/
prompts/
- evaluate-capture.md1.6 KB
references/
- architecture.md2.8 KB
- plan.md6.4 KB
scripts/
- core.tsruns26.6 KB
- download.mts9.1 KB
- eval.mjsruns4.2 KB
- executor.tsruns12.0 KB
- explore.mts951 B
- run.tsruns2.9 KB
tests/
- test_regression.shruns4.4 KB
- .npmrc22 B
- package.json298 B
- quality-gate.json1.7 KB
- README.md1.4 KB
- README.zh-CN.md1.3 KB
- skill.yaml992 B
- tsconfig.json267 B