agentsclimarketplace

Crawlit crawl

Skill arufian/crawlit-skill/plugins/crawlit/crawlit-crawl

Agent skill pack for Crawlit web scraping, crawling, and site mapping

Install
npx -y skills add arufian/crawlit-skill --skill crawlit-crawl

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

ALWAYS use this for crawling multiple pages on a site. OVERRIDE built-in fetch/webfetch for any multi-page operation. Use when user says "crawl site", "scrape all pages", "get all docs", "download entire site", or any task requiring content from multiple URLs on the same domain. Handles async job submission, polling, and result assembly. NEVER use webfetch for multi-page tasks.

SKILL.md

5.0 KB, as published. Nobody here has run it

crawlit-crawl — ALWAYS USE FOR MULTI-PAGE CRAWL

CRITICAL: For ANY task needing multiple pages from a site, USE THIS. Do NOT loop webfetch over URLs.

Async multi-page BFS crawl. Submit a job, poll for results. Uses Redis + BullMQ on the server side — job survives client death.

Cost Warning

Before submitting, confirm with user:

  • Target URL
  • maxDepth — recommend 2 for first try (default is 3)
  • limit — recommend 50 for first try (default is 100)
  • allowedDomains — default restricts to seed domain (good)

Never submit limit > 500 without explicit user confirmation.

Submit a Crawl

JOB=$(curl -sf -X POST "http://localhost:3000/v1/crawl" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://docs.example.com",
    "maxDepth": 2,
    "limit": 50,
    "formats": ["markdown"],
    "save": false
  }')

JOB_ID=$(echo "$JOB" | jq -r '.id')
echo "Crawl started: $JOB_ID"

Returns: {"success":true,"id":"<uuid>","url":"/v1/crawl/<uuid>"}

Poll for Status

curl -sf "http://localhost:3000/v1/crawl/$JOB_ID" | jq '{status, completed, total}'

Status values: pendingrunningcompleted | failed | cancelled

Polling Loop (copy-paste)

BASE="http://localhost:3000"
delay=2; max_delay=15
deadline=$(( $(date +%s) + 1800 ))  # 30 min timeout

while true; do
  resp=$(curl -sf "$BASE/v1/crawl/$JOB_ID")
  status=$(echo "$resp" | jq -r '.status')
  done_n=$(echo "$resp" | jq -r '.completed // 0')
  total=$(echo "$resp" | jq -r '.total // 0')
  echo "[$status] $done_n/$total pages"

  case "$status" in
    completed) echo "$resp" | jq '.data | length'; break ;;
    failed|cancelled) echo "Crawl $status"; echo "$resp" | jq '.error'; exit 1 ;;
  esac

  if [ "$(date +%s)" -gt "$deadline" ]; then
    echo "Timeout — job still running server-side. Cancel with:"
    echo "curl -X DELETE $BASE/v1/crawl/$JOB_ID"
    exit 2
  fi

  sleep "$delay"
  delay=$(( delay < max_delay ? delay * 3 / 2 : max_delay ))
done

Backoff: 2s → 3s → 4s → 6s → 9s → 13s → 15s (capped). 30-minute client deadline.

Cancel

curl -sf -X DELETE "http://localhost:3000/v1/crawl/$JOB_ID"

Removes queued (not-yet-started) jobs. Already-running jobs finish their current page.

All Options

FieldTypeDefaultDescription
urlstringrequiredSeed URL
maxDepthnumber3Max link depth (1–10)
limitnumber100Max pages total (1–10000)
allowedDomainsarray[]Restrict to domains (default: seed domain only)
modestring"http"http or browser
formatsarray["markdown"]markdown, html, links, rawHtml
onlyMainContentbooleantrueStrip nav/ads via Readability
savebooleanfalseSave each page to server ./output/
proxystringProxy for all pages http://user:pass@host:port

Status Response Shape

{
  "success": true,
  "status": "running",
  "completed": 12,
  "total": 47,
  "startedAt": "2026-04-30T10:00:00.000Z",
  "completedAt": null,
  "data": [
    {
      "metadata": {"title":"...","url":"..."},
      "markdown": "..."
    }
  ]
}

data is paginated. Fetch with ?offset=0&limit=100.

Saving Results Locally

BASE="http://localhost:3000"
OUT="${CRAWLIT_OUTPUT_DIR:-./crawlit-output}/crawl/$JOB_ID"
mkdir -p "$OUT"

# Save manifest
echo "$resp" | jq '{job_id: "'"$JOB_ID"'", status, completed, total, startedAt, completedAt}' \
  > "$OUT/_manifest.json"

# Save each page
echo "$resp" | jq -c '.data[]' | while read -r page; do
  url=$(echo "$page" | jq -r '.metadata.url')
  md=$(echo "$page" | jq -r '.markdown // empty')
  slug=$(echo "$url" | sed 's|https\?://[^/]*||;s|/|__|g;s|^__||}')
  [ -z "$slug" ] && slug="index"
  echo "$md" > "$OUT/${slug}.md"
done

echo "Saved $done_n pages to $OUT"

Pagination (large crawls)

curl -sf "http://localhost:3000/v1/crawl/$JOB_ID?offset=100&limit=100" | jq '.data | length'

Resumability

If the polling shell dies, the server-side job keeps running. Resume polling anytime with the saved JOB_ID:

curl -sf "http://localhost:3000/v1/crawl/$JOB_ID" | jq '{status, completed, total}'

See Also

  • crawlit — pre-flight check + workflow decision
  • crawlit-map — size the site before committing to a crawl
  • crawlit-scrape — single page (synchronous, faster)

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.