agentsclimarketplace

Pdf file organizer

Skill cxcscmu/SkillLearnBench/skills/b4-skill-creator-claude-sonnet-4-6/organize-messy-files/pdf-file-organizer

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Install
npx -y skills add cxcscmu/SkillLearnBench --skill pdf-file-organizer

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Organizes PDF, PPTX, and DOCX academic files into subject folders by extracting and analyzing their title/abstract text. Use this skill whenever the user needs to sort or categorize a batch of research papers or documents into topic-based folders (e.g., LLM, quantum computing, biology, physics, music). Handles arXiv papers and other academic documents automatically.

SKILL.md

3.5 KB, as published. Nobody here has run it

PDF File Organizer Skill

Classifies and moves academic documents (PDF, PPTX, DOCX) into subject-specific folders based on their content.

Workflow

1. Extract Text for Classification

Use pdftotext (first page only for speed) to get title + abstract:

pdftotext -l 1 file.pdf - 2>/dev/null | head -30

For PPTX/DOCX, use python-pptx or python-docx, or rely on filename/metadata.

2. Classify by Keywords

Match extracted text against subject keyword sets:

SubjectKey terms
LLMtransformer, language model, GPT, BERT, attention, token, fine-tuning, prompt, LLM, NLP, neural network, RAG, reasoning
trapped_ion_and_qctrapped ion, qubit, quantum gate, quantum circuit, quantum error, entanglement, laser cooling, quantum computing, Jaynes-Cummings
black_holeblack hole, event horizon, Hawking, gravitational wave, neutron star, spacetime, singularity, general relativity, entropy
DNADNA, gene, genome, protein, RNA, nucleotide, CRISPR, mutation, sequence, chromosome, molecular biology
music_historymusic, composer, symphony, baroque, classical period, opera, melody, harmony, instrument, jazz, folk

3. Batch Processing Strategy

For 100+ files, use a shell script to extract text and classify in bulk:

for f in /path/to/files/*.pdf; do
    text=$(pdftotext -l 1 "$f" - 2>/dev/null | head -40)
    # classify based on text content
    # move to appropriate folder
done

4. Create Target Folders

mkdir -p base_dir/{LLM,trapped_ion_and_qc,black_hole,DNA,music_history}

5. Handle Edge Cases

  • If classification is ambiguous, pick the best-matching single folder
  • Every file must end up in exactly one folder
  • Never rename files or modify their content
  • For non-PDF files (PPTX, DOCX), extract text with appropriate tools or use filename heuristics

Python Classification Script Pattern

import subprocess, shutil, os

SUBJECTS = {
    'LLM': ['language model', 'transformer', 'gpt', 'bert', 'llm', 'attention mechanism',
            'neural network', 'nlp', 'fine-tun', 'token', 'prompt', 'rag', 'reasoning model'],
    'trapped_ion_and_qc': ['trapped ion', 'qubit', 'quantum gate', 'quantum circuit',
                            'quantum error', 'entanglement', 'quantum comput', 'laser cool',
                            'jaynes-cummings', 'quant-ph', 'ion trap'],
    'black_hole': ['black hole', 'event horizon', 'hawking', 'gravitational', 'neutron star',
                   'spacetime', 'singularity', 'general relativity', 'horizon entropy'],
    'DNA': ['dna', 'gene', 'genome', 'protein', 'rna', 'nucleotide', 'crispr',
            'chromosome', 'molecular biology', 'base pair', 'sequence'],
    'music_history': ['music', 'composer', 'symphony', 'baroque', 'opera', 'melody',
                      'harmony', 'instrument', 'jazz', 'folk', 'classical music'],
}

def classify(text):
    text_lower = text.lower()
    scores = {subj: sum(1 for kw in kws if kw in text_lower)
              for subj, kws in SUBJECTS.items()}
    return max(scores, key=scores.get)

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.