Playwright scraper skill
Scrape dynamic and anti-bot protected websites using Playwright, returning page content, titles, and screenshots. Triggers when users ask to fetch, scrape, or extract data from a URL, especially for sites with JavaScript rendering, Cloudflare, or known blocking (like Discuss.com.hk).From its SKILL.md
npx -y skills add serejaris/kimi-skills --skill playwright-scraper-skillAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 5 stars5 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
6.2 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it
Playwright Scraper Skill
A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level.
π― Use Case Matrix
| Target Website | Anti-Bot Level | Recommended Method | Script |
|---|---|---|---|
| Regular Sites | Low | web_fetch tool | N/A (built-in) |
| Dynamic Sites | Medium | Playwright Simple | scripts/playwright-simple.js |
| Cloudflare Protected | High | Playwright Stealth β | scripts/playwright-stealth.js |
| YouTube | Special | deep-scraper | Install separately |
| Special | reddit-scraper | Install separately |
π¦ Installation
cd playwright-scraper-skill
npm install
npx playwright install chromium
π Quick Start
1οΈβ£ Simple Sites (No Anti-Bot)
Use OpenClaw's built-in web_fetch tool:
# Invoke directly in OpenClaw
Hey, fetch me the content from https://example.com
2οΈβ£ Dynamic Sites (Requires JavaScript)
Use Playwright Simple:
node scripts/playwright-simple.js "https://example.com"
Example output:
{
"url": "https://example.com",
"title": "Example Domain",
"content": "...",
"elapsedSeconds": "3.45"
}
3οΈβ£ Anti-Bot Protected Sites (Cloudflare etc.)
Use Playwright Stealth:
node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"
Features:
- Hide automation markers (
navigator.webdriver = false) - Realistic User-Agent (iPhone, Android)
- Random delays to mimic human behavior
- Screenshot and HTML saving support
4οΈβ£ YouTube Video Transcripts
Use deep-scraper (install separately):
# Install deep-scraper skill
npx clawhub install deep-scraper
# Use it
cd skills/deep-scraper
node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID"
π Script Descriptions
scripts/playwright-simple.js
- Use Case: Regular dynamic websites
- Speed: Fast (3-5 seconds)
- Anti-Bot: None
- Output: JSON (title, content, URL)
scripts/playwright-stealth.js β
- Use Case: Sites with Cloudflare or anti-bot protection
- Speed: Medium (5-20 seconds)
- Anti-Bot: Medium-High (hides automation, realistic UA)
- Output: JSON + Screenshot + HTML file
- Verified: 100% success on Discuss.com.hk
π Best Practices
1. Try web_fetch First
If the site doesn't have dynamic loading, use OpenClaw's web_fetch toolβit's fastest.
2. Need JavaScript? Use Playwright Simple
If you need to wait for JavaScript rendering, use playwright-simple.js.
3. Getting Blocked? Use Stealth
If you encounter 403 or Cloudflare challenges, use playwright-stealth.js.
4. Special Sites Need Specialized Skills
- YouTube β deep-scraper
- Reddit β reddit-scraper
- Twitter β bird skill
π§ Customization
All scripts support environment variables:
# Set screenshot path
SCREENSHOT_PATH=/path/to/screenshot.png node scripts/playwright-stealth.js URL
# Set wait time (milliseconds)
WAIT_TIME=10000 node scripts/playwright-simple.js URL
# Enable headful mode (show browser)
HEADLESS=false node scripts/playwright-stealth.js URL
# Save HTML
SAVE_HTML=true node scripts/playwright-stealth.js URL
# Custom User-Agent
USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js URL
π Performance Comparison
| Method | Speed | Anti-Bot | Success Rate (Discuss.com.hk) |
|---|---|---|---|
| web_fetch | β‘ Fastest | β None | 0% |
| Playwright Simple | π Fast | β οΈ Low | 20% |
| Playwright Stealth | β±οΈ Medium | β Medium | 100% β |
| Puppeteer Stealth | β±οΈ Medium | β Medium-High | ~80% |
| Crawlee (deep-scraper) | π’ Slow | β Detected | 0% |
| Chaser (Rust) | β±οΈ Medium | β Detected | 0% |
π‘οΈ Anti-Bot Techniques Summary
Lessons learned from our testing:
β Effective Anti-Bot Measures
- Hide
navigator.webdriverβ Essential - Realistic User-Agent β Use real devices (iPhone, Android)
- Mimic Human Behavior β Random delays, scrolling
- Avoid Framework Signatures β Crawlee, Selenium are easily detected
- Use
addInitScript(Playwright) β Inject before page load
β Ineffective Anti-Bot Measures
- Only changing User-Agent β Not enough
- Using high-level frameworks (Crawlee) β More easily detected
- Docker isolation β Doesn't help with Cloudflare
π Troubleshooting
Issue: 403 Forbidden
Solution: Use playwright-stealth.js
Issue: Cloudflare Challenge Page
Solution:
- Increase wait time (10-15 seconds)
- Try
headless: false(headful mode sometimes has higher success rate) - Consider using proxy IPs
Issue: Blank Page
Solution:
- Increase
waitForTimeout - Use
waitUntil: 'networkidle'or'domcontentloaded' - Check if login is required
π Memory & Experience
2026-02-07 Discuss.com.hk Test Conclusions
- β Pure Playwright + Stealth succeeded (5s, 200 OK)
- β Crawlee (deep-scraper) failed (403)
- β Chaser (Rust) failed (Cloudflare)
- β Puppeteer standard failed (403)
Best Solution: Pure Playwright + anti-bot techniques (framework-independent)
π§ Future Improvements
- Add proxy IP rotation
- Implement cookie management (maintain login state)
- Add CAPTCHA handling (2captcha / Anti-Captcha)
- Batch scraping (parallel URLs)
- Integration with OpenClaw's
browsertool
π References
What ships with it: 12 files
28.6 KB alongside SKILL.md, 4 of them executable
examples/
- discuss-hk.shruns432 B
- README.md4.1 KB
scripts/
- playwright-simple.jsruns1.7 KB
- playwright-stealth.jsruns5.4 KB
- CHANGELOG.md1.6 KB
- CONTRIBUTING.md2.9 KB
- INSTALL.md2.2 KB
- _meta.json143 B
- package.json531 B
- README.md4.4 KB
- README_ZH.md4.2 KB
- test.shruns1.1 KB