Url
Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.
npx -y skills add axoviq-ai/synthadoc --skill urlAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Fetch and extract text from web URLs
The file declares its own license as AGPL-3.0-or-later. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
1.5 KB, as published. Nobody here has run it
URL Skill
Fetches a web URL using httpx, strips navigation/script/style tags with
BeautifulSoup, and returns clean body text. PDF URLs are extracted with
pypdf (primary) and pdfminer.six (fallback).
Setup
pip install httpx beautifulsoup4
# Optional — needed only if you ingest PDF URLs:
pip install pypdf pdfminer.six
Standalone usage
import asyncio
from synthadoc.skills.url.scripts.main import UrlSkill
skill = UrlSkill()
async def main():
result = await skill.extract("https://example.com/article")
print(result.text) # clean body text
print(result.metadata) # {"url": "https://..."}
asyncio.run(main())
DomainBlockedException is raised when the site returns HTTP 401, 403, or
429. Catch it to log and skip the domain:
from synthadoc.skills.base import DomainBlockedException
try:
result = await skill.extract(url)
except DomainBlockedException as e:
print(f"Blocked: {e.domain} (HTTP {e.status_code})")
When this skill is used
- Source starts with
https://orhttp:// - User intent contains:
fetch url,web page,website