agentsclimarketplace

Hasdata

Skill HasData/agent-skills/skills/hasdata

HasData official collection of agent skills

Install
npx -y skills add HasData/agent-skills --skill hasdata

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 9 stars9 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use HasData to scrape any public web page, run real-time Google/Bing/Google-AI-Mode search queries, pull structured data from e-commerce, real-estate, lodging, jobs, maps, travel, video, and social platforms, or run async bulk-scraping and crawling jobs without managing proxies, browsers, or captchas. Reach for this skill when the user mentions web scraping, SERP/search results, Google Maps/Trends/Flights/Images, YouTube videos/transcripts/channels, Amazon, Zillow, Redfin, Airbnb, Booking.com hotels, Yelp, Indeed, Glassdoor, Instagram, Shopify, scraper jobs, website crawling, RAG/LLM data ingestion, lead/contact enrichment, or HasData itself.

SKILL.md

6.3 KB, as published. Nobody here has run it

HasData

Cloud platform for extracting public web data. One API key, three execution modes. All endpoints sit under https://api.hasdata.com and authenticate with x-api-key.

curl -G 'https://api.hasdata.com/scrape/google/serp' \
  --data-urlencode 'q=coffee' \
  -H 'x-api-key: <your-api-key>'

401 invalid key, 403 quota exhausted, 429 concurrency cap, 500 server error (retry).

Three execution modes

ModeLatencyWhenEndpoint
Web Scraping APIsecondsArbitrary URL — JS rendering, CSS/AI extraction, screenshotsPOST /scrape/web
Scraper APIs (sync)secondsPre-parsed JSON for known platforms (Google, Amazon, Zillow, …)GET /scrape/<vertical>/<resource>
Scraper Jobs (async)minutes–hoursBulk extraction, recursive crawling, webhook fan-outPOST /scrapers/<slug>/jobs

Decision rule. Default to a Scraper API when one exists for the platform (pre-parsed JSON, no selector maintenance). Use Web Scraping for arbitrary URLs not covered by an API. Reach for a Scraper Job only when no API equivalent exists — crawler, contacts, sec-edgar, amazon-bestsellers, amazon-product-reviewsor when async fan-out + webhooks save engineering time over a paginated client loop.

Always-true response shape

{ "requestMetadata": { "id": "…", "status": "ok", "url": "…" }, "...": "endpoint-specific" }

Treat data as valid only if requestMetadata.status === "ok". HTTP 200 alone isn't enough.

High-leverage patterns

  • SERP-first enrichment. Google SERP is a free-form data lake. q="<Person> <Company> linkedin"organicResults[0].title + .snippet already carries role + location, no profile-page scrape needed. Same recipe with crunchbase, wikipedia, github, or quoted literals ("[email protected]", "+1 555 …") for reverse lookup.
  • AI Mode + verify. /scrape/google/ai-mode for the answer + references → /scrape/web (markdown) on each reference URL → cited RAG context, no vector DB.
  • Maps → leads. /scrape/google-maps/search returns websites + phones; fan out to /scrape/web with extractEmails: true for full lead rows.
  • Crawler → corpus. crawler Scraper Job with outputFormat: ["markdown"] + includePaths: "/docs/.+" produces an LLM-ready corpus in one submission.
  • Pre-extracted via SERP rich snippets. knowledgeGraph, localResults, inlineShoppingResults, relatedQuestions carry pre-parsed facts that bypass anti-bot. Always check them before scraping.

When to call from code (the wiring)

  • Auth: x-api-key header on every request. Read from HASDATA_API_KEY env. Never hardcode, never log.
  • Timeouts: set client timeout ≥ 300 s. HasData's own deadline is 300 s; shorter clients produce phantom failures while still being billed on completion.
  • Retries: 429 and 5xx only — exponential backoff, jitter. Never retry 4xx (auth, validation).
  • Concurrency: cap at your plan limit. The free tier is 1; anything higher just generates 429s.
  • Async jobs: the submit response handle is body.id (integer), not jobId. Persist it immediately. Poll GET /scrapers/jobs/<id> every 10–30 s with backoff; treat webhooks as best-effort and always pair with polling. On finished the status carries data: {csv, json, xlsx} short-lived URLs — download immediately.

See references/code-recipes.md for ready-to-paste Python and TypeScript clients with retry, backoff, bounded concurrency, and the full job lifecycle.

Common gotchas

  • 300 s server deadline. Match client timeout.
  • Disable jsRendering first, enable only if the page needs it — most static pages parse fine without a headless browser.
  • No cookies parameter — cookies go through headers["Cookie"].
  • includePaths regex is case-sensitive. /blog/.+ won't match /Blog/....
  • Scraper Job data is double-wrapped. Each row is body.data[i].data; outer wraps with id, jobId, dataId, createdAt, updatedAt.
  • requestMetadata.status === "ok" is the only success signal. HTTP 200 alone isn't enough.
  • Webhooks are best-effort with 3 retries. Always have a polling fallback.

References

Resources

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.