agentsclimarketplace

Run2 robust text extraction

Skill cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/organize-messy-files/run2_robust_text_extraction

Robustly extracts text from various document formats with multiple fallback options.From its SKILL.md

Install
npx -y skills add cxcscmu/SkillLearnBench --skill run2_robust_text_extraction

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

1.2 KB, 295 tokens by cl100k_base, as published. Nobody here has run it

Robust Text Extraction Skill

Extract text from PDF, DOCX, and PPTX even when specialized tools are missing.

PDF Extraction

Primary: pdftotext -l 5 <file> - (extracts first 5 pages). Fallback: strings <file> | head -n 100 (last resort to find metadata/titles).

DOCX Extraction

Primary: pandoc <file> -t plain. Fallback (Python):

import zipfile, xml.etree.ElementTree as ET
def get_docx_text(path):
    with zipfile.ZipFile(path) as z:
        xml_content = z.read('word/document.xml')
        tree = ET.fromstring(xml_content)
        return ' '.join(node.text for node in tree.iter() if node.text)

PPTX Extraction

Primary: python3 -m markitdown <file>. Fallback (Python):

import zipfile, xml.etree.ElementTree as ET
def get_pptx_text(path):
    text = []
    with zipfile.ZipFile(path) as z:
        for f in sorted(z.namelist()):
            if f.startswith('ppt/slides/slide'):
                xml_content = z.read(f)
                tree = ET.fromstring(xml_content)
                text.append(' '.join(node.text for node in tree.iter() if node.text))
    return ' '.join(text)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.