Run1 enterprise data parsing
Strategies for reading and indexing diverse enterprise data formats located in a flat directory or specific path.From its SKILL.md
npx -y skills add cxcscmu/SkillLearnBench --skill run1_enterprise-data-parsingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
0.9 KB, 169 tokens by cl100k_base, as published. Nobody here has run it
When retrieving information from enterprise data (like /root/DATA), first identify the file structure. Enterprise data often consists of mixed formats (PDF, CSV, JSON, TXT, or Markdown).
- Inventory: List all files in the directory to determine the scope.
import os data_path = "/root/DATA" files = os.listdir(data_path) - Reading Techniques:
- Text/Markdown: Use standard
open().read(). - JSON: Use
json.load(). - CSV: Use
pandas.read_csv()for structured queries.
- Text/Markdown: Use standard
- Indexing: If the dataset is large, create a simple keyword index or use a search function to map keywords from the questions to specific filenames to narrow the search space.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.