The catalog
325,949 rows, by stacks then stars
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_csv-reporting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_cvss-score Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_trivy-offline Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_antimeridian-handling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 geopandas geospatial distance
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_geopandas-geospatial-distance Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_plate-boundary-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 enterprise data retrieval
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/enterprise-information-search/run2_enterprise-data-retrieval Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/enterprise-information-search/run2_slack-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/financial-analysis/run2_13f-data-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_jackson-inject-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/github-repo-analytics/run2_data-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/github-repo-analytics/run2_gh-rest-api Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 pytorch preference optimization
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/nlp-paper-reproduction/run2_pytorch-preference-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/nlp-paper-reproduction/run2_simpo-loss Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 python docx split placeholders
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/offer-letter-generator/run2_python-docx-split-placeholders Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/organize-messy-files/run2_document-text-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/organize-messy-files/run2_keyword-classification Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-circe-json Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-idiomatic-patterns Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-immutable-builder Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-option-handling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-variance Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_meeting-scheduling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_pdf-calendar-parsing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_timezone-conversion Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/stock-data-visualization/run2_d3-force-bubble Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/stock-data-visualization/run2_d3-interactive-table Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_glm-calibration-strategy Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_glm-lake-modeling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_netcdf-obs-matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_budget-validation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_itinerary-planning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_travel-data-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_ffmpeg-keyframes Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_opencv-grayscale Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_template-matching-tuning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/weighted-gdp-calculation/run2_excel-index-match Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/weighted-gdp-calculation/run2_excel-weighted-mean Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_anthropic-brand-tokens Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_exploded-view-drawing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 matplotlib technical poster
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_matplotlib-technical-poster Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 chinese poetry composition
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/chinese-poem-generator/run2_chinese-poetry-composition Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/chinese-poem-generator/run2_qiyan-lushi Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/court-form-filling/run2_sc100-small-claims-form Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_dbscan-clustering Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_greedy-matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_pareto-frontier Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_cvss-score-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_trivy-offline-scanning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 vulnerability csv reporting
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_vulnerability-csv-reporting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run2_earthquake-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 geopandas plate tectonics
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run2_geopandas-plate-tectonics Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/enterprise-information-search/run2_competitor-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 enterprise data retrieval
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/enterprise-information-search/run2_enterprise-data-retrieval Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_13f-fund-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_fuzzy-fund-search Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_holdings-comparison Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_palantir-holders Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.