agentsclimarketplace

Crawl

Skill eprouveze/claude-skills/skills/crawl

A collection of practical Claude Code skills — multi-LLM evaluation, domain management, planning, writing quality, and pair-session patterns. MIT licensed.

Install
npx -y skills add eprouveze/claude-skills --skill crawl

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Fetch web pages that may be JS-rendered or bot-protected, returning clean markdown or HTML. Works zero-setup with a direct HTTP fetch; if you supply your own Cloudflare Browser Rendering credentials it uses a managed headless browser that renders JS and bypasses most WAFs. Use when a plain fetch returns 403, when a page is a JS-rendered SPA, or when you need reliable markdown extraction from a URL. Triggers on 'crawl this page', 'fetch this URL', 'scrape this site', 'get the content from this page', 'this page is blocked', or when a normal fetch fails on a URL.

SKILL.md

7.8 KB, as published. Nobody here has run it

/crawl — Tiered Web Page Fetcher

Fetch web pages that may be JS-rendered or bot-protected, and return clean markdown (or raw HTML). It has two tiers: a direct HTTP fetch that needs zero configuration, and Cloudflare Browser Rendering — a managed headless Chromium that renders JavaScript, follows redirects, and bypasses most WAFs — which activates only when you supply your own Cloudflare credentials.

The bundled script is scripts/crawl.ts (run with npx tsx).

Requirements

Out of the box (no setup): the direct-fetch tier needs only Node.js and tsx (npm install -g tsx). It handles static HTML, many simple pages, and .md/.md.txt/.txt endpoints. JS-rendered SPAs and bot-protected pages will fail this tier — that's expected, and where the Cloudflare tier comes in.

Optional — bring your own Cloudflare Browser Rendering (recommended for JS/SPA/blocked pages): Cloudflare's Browser Rendering REST API runs a real headless Chromium in their network. It's a per-account API — there's no shared endpoint, so you supply your own account ID and token via environment variables:

VariableRequired forWhat it is
CLOUDFLARE_ACCOUNT_IDCF tierYour Cloudflare account ID (32-char hex)
CLOUDFLARE_BR_TOKENCF tierAn API token scoped to Account → Browser Rendering → Edit

Setup:

  1. Sign in at https://dash.cloudflare.com and copy your Account ID (right sidebar).
  2. Create an API token at https://dash.cloudflare.com/profile/api-tokens with the Account → Browser Rendering → Edit permission. Copy it once — it isn't shown again.
  3. Run the interactive setup, which writes ~/.config/claude-skills/crawl.env (mode 600) and is auto-loaded on every run:
    npx tsx scripts/crawl.ts --setup
    
    Prefer env vars? Export them instead (or put them in ~/.env) — both are auto-loaded, with the skill config taking precedence:
    export CLOUDFLARE_ACCOUNT_ID="your-account-id"
    export CLOUDFLARE_BR_TOKEN="your-api-token"
    

Cloudflare Browser Rendering docs: https://developers.cloudflare.com/browser-rendering/rest-api/ The free plan includes a daily quota; higher volume needs a Workers Paid plan.

If the Cloudflare variables are unset, the CF tier is skipped automatically and the skill operates in direct-fetch-only mode.

Usage

From the CLI

# Single URL → markdown to stdout
npx tsx scripts/crawl.ts "https://example.com"

# Save to file (writes YAML frontmatter + content)
npx tsx scripts/crawl.ts "https://example.com" --save ./output.md

# Get HTML instead of markdown
npx tsx scripts/crawl.ts "https://example.com" --format html

# Force a tier: cf = Cloudflare only, direct = no CF (e.g. a known plain-text endpoint)
npx tsx scripts/crawl.ts "https://example.com" --tier cf
npx tsx scripts/crawl.ts "https://example.com/llms.txt" --tier direct

# Batch mode (one URL per line; output dir for --save)
npx tsx scripts/crawl.ts --batch urls.txt --save ./output/

# JSON output (for scripting)
npx tsx scripts/crawl.ts "https://example.com" --json

# Add a short LLM-generated summary (needs GEMINI_API_KEY — see below)
npx tsx scripts/crawl.ts "https://example.com" --summarize

From TypeScript (import into another script)

import { crawl } from "./scripts/crawl.ts";

const result = await crawl("https://example.com", { format: "markdown" });
if (result.content && !result.error) {
  // result.content is the markdown string
} else {
  // result.error / result.blocked tell you what went wrong
}

Options

FlagValuesDefaultNotes
--formatmarkdown | htmlmarkdownOutput format
--tierall | cf | directallWhich tiers to try
--savepathFile (single) or directory (batch)
--batchfileOne URL per line; # lines ignored
--jsonoffEmit a JSON CrawlResult instead of raw content
--summarizeoffAppend a 2–3 sentence LLM summary

How it works

Two tiers. With the default --tier all:

  1. Cloudflare Browser Rendering is tried first when credentials are set — it renders JS, bypasses most WAFs, follows redirects, and returns markdown or HTML. If the CLOUDFLARE_* vars are missing, this tier returns immediately with a "not set" error and the skill falls through to tier 2.
  2. Direct fetch (always available) — a plain HTTP fetch with a desktop User-Agent, basic HTML→text conversion, and an automatic HTTPS→HTTP retry for legacy sites. It detects WAF block pages and unrendered "Loading…" SPAs and reports them as errors.

So out of the box (no Cloudflare account) every call degrades gracefully to direct fetch. Use --tier direct to skip CF entirely (e.g. a known llms.txt/.md.txt endpoint), or --tier cf to skip the direct attempt.

If both tiers fail, the result carries error (and blocked: true for WAF pages) so the caller can pick the next move — an official docs API, a GitHub-hosted README, a search tool.

What each tier handles

ScenarioDirectCloudflare Browser Rendering
Static HTMLworksworks
JS SPA (React/Vue)"Loading…" onlyworks (renders JS)
Bot-protected pagesoften 403usually works
.md.txt / llms.txt endpointsworksworks
Sites that block Cloudflare's egress IPsvariesblocked (WAF blocks CF too)

Optional summarization

--summarize appends a short, fact-focused summary generated by Google's Gemini Flash. It needs a GEMINI_API_KEY (or GOOGLE_AI_API_KEY) in the environment. If neither is set, summarization is silently skipped and crawling still works. Swap in any LLM you prefer by editing summarizeContent() in scripts/crawl.ts.

Output format

With --save, the script writes a YAML frontmatter block followed by the content:

---
url: https://example.com
tier: cf-markdown
fetched_at: 2026-06-06T12:00:00.000Z
content_length: 4821
http_status: 200
---

error, blocked, and summary fields are added when present. Without --save, content goes to stdout and diagnostics go to stderr, so you can pipe cleanly.

Anti-patterns

  • Don't hardcode credentials in scripts or skill files. The Cloudflare token belongs in an environment variable.
  • Don't loop-crawl hundreds of URLs without delays — you'll hit Cloudflare's quota. Use --batch (capped at concurrency 5) and space out large jobs.
  • Don't assume CF can reach everything. Some origins block Cloudflare's IPs; keep a non-CF fallback ready.
  • Don't use --tier cf with no credentials — it just errors. Let the default all degrade to direct fetch.

Error handling

  • error set + empty content → the tier failed; the message says why (HTTP 403, Timeout, CLOUDFLARE_BR_TOKEN ... not set, etc.).
  • blocked: true → a WAF/CAPTCHA block page was detected. Switch tiers or use another source.
  • Exit codes: 0 success, 1 all tiers failed, 2 invalid arguments.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.