The catalog
325,949 rows, by stacks then stars
Run2 spatial analysis validation
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/earthquake-plate-calculation/run2_spatial-analysis-validation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/enterprise-information-search/run2_advanced-slack-mining Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 document metadata extraction
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/enterprise-information-search/run2_document-metadata-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 multi source data correlation
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/enterprise-information-search/run2_multi-source-data-correlation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/financial-analysis/run2_cross_fund_analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/financial-analysis/run2_fund_analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/financial-analysis/run2_quarter_comparison Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/financial-analysis/run2_sec13f_data_loader Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 druid javascript deserialization security
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_druid-javascript-deserialization-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 jackson security patching
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_jackson-security-patching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_maven-security-build Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/github-repo-analytics/run2_data-aggregation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/github-repo-analytics/run2_github-rest-api Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/github-repo-analytics/run2_json-validation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 simpo loss detailed implementation
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/nlp-paper-reproduction/run2_simpo-loss-detailed-implementation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 simpo testing and verification
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/nlp-paper-reproduction/run2_simpo-testing-and-verification Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 simpo trainer initialization
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/nlp-paper-reproduction/run2_simpo-trainer-initialization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/offer-letter-generator/run2_placeholder-replacement Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/offer-letter-generator/run2_python-docx-basics Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 content classification enhanced
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/organize-messy-files/run2_content-classification-enhanced Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/organize-messy-files/run2_file-organization-robust Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/organize-messy-files/run2_multi-format-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/organize-messy-files/run2_pdf-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/python-scala-translation/run2_scala-immutable-builder Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/python-scala-translation/run2_scala-java-time-handling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 scala pattern matching dispatch
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/python-scala-translation/run2_scala-pattern-matching-dispatch Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/python-scala-translation/run2_scala-string-processing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/schedule-planning/run2_pdf-visual-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/schedule-planning/run2_request-parsing-advanced Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/schedule-planning/run2_time-formatting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/stock-data-visualization/run2_csv_robust_loading Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/stock-data-visualization/run2_d3_bubble_clusters Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/stock-data-visualization/run2_d3_sync_interaction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/stock-data-visualization/run2_market_cap_formatting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_glm-error-diagnosis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_netcdf-glm-matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_nml-calibration-workflow Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 parameter optimization strategy
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_parameter-optimization-strategy Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 complete itinerary solution
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/travel-planning/run2_complete_itinerary_solution Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 data filtering validation
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/travel-planning/run2_data_filtering_validation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/travel-planning/run2_itinerary_optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 advanced template matching
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/video-object-counting/run2_advanced_template_matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/video-object-counting/run2_csv_validation_export Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/video-object-counting/run2_ffmpeg_robust_extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 image processing verification
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/video-object-counting/run2_image_processing_verification Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/weighted-gdp-calculation/run2_excel-percentages Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 excel statistics aggregates
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/weighted-gdp-calculation/run2_excel-statistics-aggregates Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/weighted-gdp-calculation/run2_multi-condition-lookups Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/weighted-gdp-calculation/run2_openpyxl-cross-sheet Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/anthropic-poster-design/run2_anthropic-brand Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/anthropic-poster-design/run2_exploded-view-layout Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/anthropic-poster-design/run2_pillow-technical-poster Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/chinese-poem-generator/run2_mandarin-rhyming Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/chinese-poem-generator/run2_peace-poem-composition Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 seven char regulated verse
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/chinese-poem-generator/run2_seven-char-regulated-verse Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/court-form-filling/run2_ca-small-claims Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/court-form-filling/run2_pdf-form-filling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dbscan-parameter-tuning/run2_dbscan-custom-metric Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dbscan-parameter-tuning/run2_parallel-grid-search Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dbscan-parameter-tuning/run2_pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.