Run2 pdf text extraction
Extract text from PDF files using PyPDF2 with pdfplumber fallback; optimized for title/abstract extraction for classification.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run2_pdf-text-extractionAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
2.0 KB, 397 tokens by cl100k_base, as published. Nobody here has run it
PDF Text Extraction Skill (Improved)
Overview
Extract text from PDF files for content analysis. The first 3-4 pages contain title, abstract, and introduction — sufficient for classification.
Installation
pip install PyPDF2 pdfplumber --break-system-packages
Key Insight
- PyPDF2 is faster but may miss some text
- pdfplumber handles complex layouts better
- For classification, extracting title + abstract (first 2 pages) is usually sufficient
- Always use both libraries with fallback for robustness
Robust Extraction Function
import PyPDF2
import pdfplumber
def extract_pdf_text(filepath, max_pages=4):
"""Extract text from PDF, trying PyPDF2 first, then pdfplumber fallback."""
text = ""
try:
with open(filepath, 'rb') as f:
reader = PyPDF2.PdfReader(f)
for page in reader.pages[:max_pages]:
try:
page_text = page.extract_text()
if page_text:
text += page_text + "\n"
except:
pass
except Exception:
pass
# If PyPDF2 didn't get much, try pdfplumber
if len(text.strip()) < 200:
try:
with pdfplumber.open(filepath) as pdf:
for page in pdf.pages[:max_pages]:
try:
page_text = page.extract_text()
if page_text:
text += page_text + "\n"
except:
pass
except Exception:
pass
return text
Notes
- Some PDFs with score=0 may need manual review — check the title from the first few lines
- arXiv PDFs generally work well with both libraries
- PDFs with equations may have garbled text but keywords like "trapped ion", "DNA" still appear
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.