agentsclimarketplace

Harvest deep crawl

Skill vibeeval/vibecosystem/skills/harvest-deep-crawl

AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution.

Install
npx -y skills add vibeeval/vibecosystem --skill harvest-deep-crawl

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Multi-page deep crawling - documentation sites, wikis, knowledge bases

SKILL.md

2.5 KB, as published. Nobody here has run it

Harvest Deep Crawl

Crawl multi-page websites following internal links to a specified depth. Ideal for building complete knowledge bases from documentation sites, wikis, and reference materials.

Usage

/crawl <url> --depth <N>

Examples

# Crawl docs site 3 levels deep
/crawl https://docs.example.com --depth 3

# Crawl a specific section
/crawl https://docs.example.com/api --depth 2

# Crawl with page limit
/crawl https://wiki.example.com --depth 5 --max-pages 50

Parameters

ParamDefaultDescription
--depth2Max link-following depth
--max-pages100Max pages to crawl
--same-domaintrueStay on same domain
--include*URL pattern to include
--exclude-URL pattern to exclude

How It Works

  1. Start at root URL, extract all internal links
  2. Follow links up to specified depth (BFS order)
  3. Extract content from each page
  4. Deduplicate pages with > 90% content overlap
  5. Build table of contents from page hierarchy
  6. Merge into coherent knowledge base
  7. Save to .claude/cache/agents/harvest/crawl-{domain}/

Output Structure

crawl-{domain}-{timestamp}/
  index.md          # Table of contents + summary
  page-001.md       # First page content
  page-002.md       # Second page content
  ...
  metadata.json     # Crawl stats, URLs, timings

Crawl Engine

Primary: crawl4ai (Docker port 11235)

curl -s http://localhost:11235/crawl \
  -H "Content-Type: application/json" \
  -d '{
    "urls": ["https://docs.example.com"],
    "max_depth": 3,
    "same_domain": true,
    "word_count_threshold": 50
  }'

Fallback: Manual Link Following

When Docker unavailable:

  1. WebFetch root URL
  2. Parse links from markdown output
  3. WebFetch each linked page (depth-limited)
  4. Compile results

Use Cases

ScenarioDepthMax Pages
API reference2-350
Full documentation site3-5100
Wiki section230
Changelog history1-220
Tutorial series2-330

Rules

  • Respect robots.txt
  • Max 2 requests/second
  • Skip binary files (PDF, images, videos)
  • Detect and skip infinite pagination
  • Cache results for 24 hours

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.