agentsclimarketplace

Run2 pdf text extraction

Skill cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/organize-messy-files/run2_pdf-text-extraction

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run2_pdf-text-extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Extract text from PDF files using PyPDF2 with pdfplumber fallback; optimized for title/abstract extraction for classification.

SKILL.md

2.0 KB, as published. Nobody here has run it

PDF Text Extraction Skill (Improved)

Overview

Extract text from PDF files for content analysis. The first 3-4 pages contain title, abstract, and introduction — sufficient for classification.

Installation

pip install PyPDF2 pdfplumber --break-system-packages

Key Insight

  • PyPDF2 is faster but may miss some text
  • pdfplumber handles complex layouts better
  • For classification, extracting title + abstract (first 2 pages) is usually sufficient
  • Always use both libraries with fallback for robustness

Robust Extraction Function

import PyPDF2
import pdfplumber

def extract_pdf_text(filepath, max_pages=4):
    """Extract text from PDF, trying PyPDF2 first, then pdfplumber fallback."""
    text = ""
    try:
        with open(filepath, 'rb') as f:
            reader = PyPDF2.PdfReader(f)
            for page in reader.pages[:max_pages]:
                try:
                    page_text = page.extract_text()
                    if page_text:
                        text += page_text + "\n"
                except:
                    pass
    except Exception:
        pass

    # If PyPDF2 didn't get much, try pdfplumber
    if len(text.strip()) < 200:
        try:
            with pdfplumber.open(filepath) as pdf:
                for page in pdf.pages[:max_pages]:
                    try:
                        page_text = page.extract_text()
                        if page_text:
                            text += page_text + "\n"
                    except:
                        pass
        except Exception:
            pass

    return text

Notes

  • Some PDFs with score=0 may need manual review — check the title from the first few lines
  • arXiv PDFs generally work well with both libraries
  • PDFs with equations may have garbled text but keywords like "trapped ion", "DNA" still appear

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.