agentsclimarketplace

Pdf file organizer

Skill cxcscmu/SkillLearnBench/skills/b4-skill-creator-claude-sonnet-4-6/organize-messy-files/pdf-file-organizer

Organizes PDF, PPTX, and DOCX academic files into subject folders by extracting and analyzing their title/abstract text. Use this skill whenever the user needs to sort or categorize a batch of research papers or documents into topic-based folders (e.g., LLM, quantum computing, biology, physics, music). Handles arXiv papers and other academic documents automatically.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill pdf-file-organizer

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

3.5 KB, 841 tokens by cl100k_base, as published. Nobody here has run it

PDF File Organizer Skill

Classifies and moves academic documents (PDF, PPTX, DOCX) into subject-specific folders based on their content.

Workflow

1. Extract Text for Classification

Use pdftotext (first page only for speed) to get title + abstract:

pdftotext -l 1 file.pdf - 2>/dev/null | head -30

For PPTX/DOCX, use python-pptx or python-docx, or rely on filename/metadata.

2. Classify by Keywords

Match extracted text against subject keyword sets:

SubjectKey terms
LLMtransformer, language model, GPT, BERT, attention, token, fine-tuning, prompt, LLM, NLP, neural network, RAG, reasoning
trapped_ion_and_qctrapped ion, qubit, quantum gate, quantum circuit, quantum error, entanglement, laser cooling, quantum computing, Jaynes-Cummings
black_holeblack hole, event horizon, Hawking, gravitational wave, neutron star, spacetime, singularity, general relativity, entropy
DNADNA, gene, genome, protein, RNA, nucleotide, CRISPR, mutation, sequence, chromosome, molecular biology
music_historymusic, composer, symphony, baroque, classical period, opera, melody, harmony, instrument, jazz, folk

3. Batch Processing Strategy

For 100+ files, use a shell script to extract text and classify in bulk:

for f in /path/to/files/*.pdf; do
    text=$(pdftotext -l 1 "$f" - 2>/dev/null | head -40)
    # classify based on text content
    # move to appropriate folder
done

4. Create Target Folders

mkdir -p base_dir/{LLM,trapped_ion_and_qc,black_hole,DNA,music_history}

5. Handle Edge Cases

  • If classification is ambiguous, pick the best-matching single folder
  • Every file must end up in exactly one folder
  • Never rename files or modify their content
  • For non-PDF files (PPTX, DOCX), extract text with appropriate tools or use filename heuristics

Python Classification Script Pattern

import subprocess, shutil, os

SUBJECTS = {
    'LLM': ['language model', 'transformer', 'gpt', 'bert', 'llm', 'attention mechanism',
            'neural network', 'nlp', 'fine-tun', 'token', 'prompt', 'rag', 'reasoning model'],
    'trapped_ion_and_qc': ['trapped ion', 'qubit', 'quantum gate', 'quantum circuit',
                            'quantum error', 'entanglement', 'quantum comput', 'laser cool',
                            'jaynes-cummings', 'quant-ph', 'ion trap'],
    'black_hole': ['black hole', 'event horizon', 'hawking', 'gravitational', 'neutron star',
                   'spacetime', 'singularity', 'general relativity', 'horizon entropy'],
    'DNA': ['dna', 'gene', 'genome', 'protein', 'rna', 'nucleotide', 'crispr',
            'chromosome', 'molecular biology', 'base pair', 'sequence'],
    'music_history': ['music', 'composer', 'symphony', 'baroque', 'opera', 'melody',
                      'harmony', 'instrument', 'jazz', 'folk', 'classical music'],
}

def classify(text):
    text_lower = text.lower()
    scores = {subj: sum(1 for kw in kws if kw in text_lower)
              for subj, kws in SUBJECTS.items()}
    return max(scores, key=scores.get)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.