agentsclimarketplace

Web scraping

Skill pencil20388-eng/browser-automation-skills/skills/web-scraping

🌐 Drop-in skills that teach Claude Code, Cursor & Codex CLI how to automate browsers. Playwright, Selenium, Puppeteer, Web Scraping, AdsPower & more.From the repository description

Install
npx -y skills add pencil20388-eng/browser-automation-skills --skill web-scraping

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
  • runs commandsInstructs the agent to run 2 commands, including `pip install requests beautifulsoup4 lxml` and 1 more.

SKILL.md

4.2 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it

Web Scraping Skill

Extract structured data from websites. Covers static scraping (fast, lightweight) and dynamic scraping (for JavaScript-rendered pages).

When to use what

ApproachWhen to useSpeedLibraries
Static (requests + BS4)Server-rendered HTML, APIsFastrequests, beautifulsoup4, lxml
Dynamic (browser)JavaScript-rendered, SPAsSlowerplaywright, selenium

Default to static scraping. Only use browser-based scraping if the content is rendered by JavaScript.

Installation

# Static scraping
pip install requests beautifulsoup4 lxml

# Dynamic scraping (if needed)
pip install playwright && playwright install chromium

Static Scraping Patterns

Basic page fetch + parse

import requests
from bs4 import BeautifulSoup

resp = requests.get("https://example.com", timeout=10)
resp.raise_for_status()
soup = BeautifulSoup(resp.text, "lxml")

title = soup.find("h1").text
links = [a["href"] for a in soup.select("a[href]")]

Extract structured data

products = []
for card in soup.select(".product-card"):
    products.append({
        "name": card.select_one(".name").text.strip(),
        "price": card.select_one(".price").text.strip(),
        "url": card.select_one("a")["href"],
    })

Handle pagination

all_items = []
page = 1

while True:
    resp = requests.get(f"https://example.com/products?page={page}", timeout=10)
    soup = BeautifulSoup(resp.text, "lxml")

    items = soup.select(".product-card")
    if not items:
        break  # No more pages

    for item in items:
        all_items.append(item.select_one(".name").text.strip())

    page += 1
    time.sleep(1)  # Be polite

print(f"Scraped {len(all_items)} items across {page - 1} pages")

Set headers to avoid blocks

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
}
resp = requests.get(url, headers=headers, timeout=10)

Handle cookies and sessions

session = requests.Session()
# Login first
session.post("https://example.com/login", data={"user": "x", "pass": "y"})
# Then scrape authenticated pages
resp = session.get("https://example.com/dashboard")

Save to CSV

import csv

with open("output.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "price", "url"])
    writer.writeheader()
    writer.writerows(products)

Dynamic Scraping (JavaScript pages)

Use Playwright when the page content is rendered by JavaScript:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://spa-example.com")
    page.wait_for_selector(".product-card")

    items = page.query_selector_all(".product-card")
    for item in items:
        name = item.text_content()
        print(name)

    browser.close()

Anti-bot handling

TechniqueCounter-measure
Rate limitingAdd time.sleep(1-3) between requests
User-Agent checkSet a realistic User-Agent header
Cookie/session checkUse requests.Session()
JavaScript renderingUse Playwright or Selenium
CAPTCHAConsider using a CAPTCHA solving service, or scrape a different source
IP blockingRotate proxies

CSS Selector cheat sheet

SelectorMeaning
div.classnamediv with class
#idnameElement with ID
div > pDirect child
div pAny descendant
a[href]Element with attribute
a[href*="keyword"]Attribute contains
.card:nth-child(2)Second child
div.card + div.cardAdjacent sibling

Important notes

  • Always set a timeout on requests: requests.get(url, timeout=10)
  • Respect robots.txt and rate limit your requests
  • Use lxml parser over html.parser β€” it's faster and handles broken HTML better
  • Check if the site has a public API first β€” it's always better than scraping
  • Store raw HTML before parsing; if your parser fails, you don't need to re-fetch

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.