agentsclimarketplace

Facebook scraper

Skill chenwei791129/agent-skills/skills/facebook-scraper

My agent skills

Install
npx -y skills add chenwei791129/agent-skills --skill facebook-scraper

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Scrape posts (text, images, top-level comments) from a Facebook group or fan page and output a Markdown report. Use when: the user wants to scrape, back up, or collect FB group/page post content and comments. Trigger on: facebook group, fan page, post scrape, fb group/page scraper.

SKILL.md

2.5 KB, as published. Nobody here has run it

Facebook Scraper

Scrape posts from a Facebook group or fan page using patchright (anti-detection Playwright): author, time, text, images, and top-level comments, rendered as a Markdown report.

Prerequisites

  1. Install Chrome for patchright (one-time):
    uv run --with patchright patchright install chrome
    
  2. Provide credentials (either option):
    • Environment variables FB_EMAIL / FB_PASSWORD
    • A .env file in the current directory or the skill directory (copy .env.example and fill it in)

Usage

uv run ~/.claude/skills/facebook-scraper/scripts/fb_scraper.py \
  --url <group-or-page-URL> [--limit N] [--limit-hour H] [--out DIR]
FlagDefaultDescription
--urlrequiredGroup or fan page URL
--limit10Number of most-recent posts
--limit-hournoneOnly posts within the last H hours (intersected with --limit)
--out/tmp/fb_scrape_output/<slug>_<timestamp>/Output directory

When both --limit and --limit-hour are given they are intersected (no more than N posts AND no older than H hours; scrolling stops as soon as either bound is hit).

On first run a browser opens: if not logged in it logs in automatically with the credentials; if 2FA / a security checkpoint appears it pauses and asks you to finish it in the open browser, then press Enter. The session is saved under .userdata/, so later runs usually skip login.

Output

  • report.md: author / time / permalink / text / images / top-level comments for each post
  • images/: downloaded images

How it works

Two phases: first scroll the feed to collect each post's permalink (applying limit / limit-hour), then visit each permalink to fetch the full comment thread and resolve the authoritative author. Post content (author, time, text, images) is extracted from the feed, where each post container is cleanly separated. All DOM extraction logic lives in scripts/extract.js; selector references and maintenance notes are in references/selectors.md (check there first when FB changes its markup).

Limitations (non-goals)

No posting, commenting, or liking.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.