agentsclimarketplace

Pdf

Skill axoviq-ai/synthadoc/synthadoc/skills/pdf

Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.

Install
npx -y skills add axoviq-ai/synthadoc --skill pdf

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Extract text from PDF documents

The file declares its own license as AGPL-3.0-or-later. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

1.1 KB, as published. Nobody here has run it

PDF Skill

Extracts text from PDF files using pypdf as the primary parser, with pdfminer.six as a fallback for CJK fonts that pypdf cannot decode (detected when pypdf yields fewer than 50 characters per page on average).

Setup

pip install pypdf pdfminer.six

Standalone usage

import asyncio
from synthadoc.skills.pdf.scripts.main import PdfSkill

skill = PdfSkill()

async def main():
    result = await skill.extract("/path/to/paper.pdf")
    print(result.text)          # extracted text from all pages
    print(result.metadata)      # {"pages": N, "cjk_fallback": bool, ...}

asyncio.run(main())

When this skill is used

  • Source path ends with .pdf
  • User intent contains: pdf, research paper

Scripts

  • scripts/main.pyPdfSkill class

References

  • references/cjk-notes.md — notes on CJK font handling

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.