Defuddle
Strip clutter from web pages before ingesting into the wiki. Removes ads, navigation, headers, footers, and boilerplate: leaving clean readable markdown that saves 40-60% tokens. Triggers on: defuddle, clean this page, strip this url, fetch and clean, clean web content before ingesting, strip ads, remove clutter, clean URL content, readable markdown from URL.From its SKILL.md
npx -y skills add eliransu/digital-brain --skill defuddleAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
3 things to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 7 commands, including `npm install -g defuddle-cli` and 6 more.
- fetches URLsInstructs the agent to fetch 1 URL, including https://example.com/article.
SKILL.md
2.6 KB, 541 tokens by cl100k_base, as published. Nobody here has run it
defuddle: Web Page Cleaner
Defuddle extracts the meaningful content from a web page and drops everything else: ads, cookie banners, nav bars, related articles, footers, social sharing buttons. What remains is the article body as clean markdown.
Use this before any URL ingestion. It is optional but strongly recommended. It cuts token usage by 40-60% on typical web articles and produces cleaner wiki pages.
Install
npm install -g defuddle-cli
Verify: defuddle --version
Usage
Clean a URL directly
defuddle https://example.com/article
Outputs clean markdown to stdout.
Save to .raw/
defuddle https://example.com/article > .raw/articles/article-slug-$(date +%Y-%m-%d).md
Add frontmatter header after saving
After running defuddle, prepend the source URL and fetch date:
SLUG="article-slug-$(date +%Y-%m-%d)"
{ echo "---"; echo "source_url: https://example.com/article"; echo "fetched: $(date +%Y-%m-%d)"; echo "---"; echo ""; defuddle https://example.com/article; } > .raw/articles/$SLUG.md
Clean a local HTML file
defuddle page.html
When to Use
Use defuddle when:
- Ingesting a news article, blog post, or documentation page from a URL
- The page has a lot of surrounding content (most web pages do)
- You want to stay within token budget on a long article
Skip defuddle when:
- The source is already a clean markdown or PDF file
- The page is a dashboard, app, or structured data (defuddle expects article-style content)
- defuddle is not installed and the article is short enough to process raw
Fallback
If defuddle is not installed, check:
which defuddle 2>/dev/null || echo "not installed"
If not installed: use WebFetch directly. The content will be less clean but still workable.
Integration with /wiki-ingest
The /wiki-ingest skill checks for defuddle automatically when a URL is passed. You do not need to run defuddle manually before ingesting a URL. The ingest skill will call it if available.
To manually clean a page and save before ingesting:
- Run the save command above
- Then:
ingest .raw/articles/[slug].md
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.