The catalog
325,949 rows, by stacks then stars
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/fix-security-bug/run2_jackson-inject-bypass Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/github-repo-analytics/run2_gh-issue-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/github-repo-analytics/run2_gh-pr-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/github-repo-analytics/run2_gh-search-api Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/github-repo-analytics/run2_report-builder Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run2_nlp-paper-reproduction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run2_simpo-loss Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 docx conditional sections
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/offer-letter-generator/run2_docx-conditional-sections Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/offer-letter-generator/run2_python-docx-placeholders Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 file classification improved
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/organize-messy-files/run2_file-classification-improved Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/organize-messy-files/run2_pdf-text-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/organize-messy-files/run2_pptx-docx-reading Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/python-scala-translation/run2_circe-json Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/python-scala-translation/run2_scala-builder Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/python-scala-translation/run2_scala-sealed-traits Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/python-scala-translation/run2_scala-type-classes Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/python-scala-translation/run2_scala-variance Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 meeting scheduler with output
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/schedule-planning/run2_meeting-scheduler-with-output Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/schedule-planning/run2_pdf-calendar-parsing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/schedule-planning/run2_timezone-dst Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/stock-data-visualization/run2_d3-bubble-chart Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/stock-data-visualization/run2_d3-force-cluster Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/stock-data-visualization/run2_d3-interactive-table Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/temperature-simulation/run2_glm-calibration Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/temperature-simulation/run2_glm-netcdf-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/temperature-simulation/run2_glm-rmse-metrics Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 itinerary data exploration
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/travel-planning/run2_itinerary-data-exploration Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/travel-planning/run2_itinerary-planning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/video-object-counting/run2_ffmpeg-keyframes Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/video-object-counting/run2_image-grayscale Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/video-object-counting/run2_object-counting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/weighted-gdp-calculation/run2_excel-index-match Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/weighted-gdp-calculation/run2_excel-sumproduct Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/weighted-gdp-calculation/run2_openpyxl-editing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 advanced technical pillow
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/anthropic-poster-design/run2_advanced-technical-pillow Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/anthropic-poster-design/run2_anthropic-brand-identity Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/chinese-poem-generator/run2_classical-poetry-mastery Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/chinese-poem-generator/run2_literary-emotive-writing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 mandarin rhyme comprehensive
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/chinese-poem-generator/run2_mandarin-rhyme-comprehensive Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/court-form-filling/run2_pdf-field-mapping Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 pdf form filling advanced
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/court-form-filling/run2_pdf-form-filling-advanced Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run2_advanced-pareto Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run2_matching-metrics Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run2_optimized-dbscan Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run2_parallel-orchestration Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dependency-vulnerability-check/run2_csv-reporting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dependency-vulnerability-check/run2_trivy-offline Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dependency-vulnerability-check/run2_vulnerability-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/earthquake-plate-calculation/run2_advanced_projections Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/earthquake-plate-calculation/run2_antimeridian_handling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 enterprise product data analysis
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/enterprise-information-search/run2_enterprise-product-data-analysis Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 user identity feedback extraction
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/enterprise-information-search/run2_user-identity-feedback-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/financial-analysis/run2_fund-analysis-pro Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/financial-analysis/run2_fund-search-pro Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/financial-analysis/run2_stock-analysis-pro Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/fix-security-bug/run2_druid-build-maven Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/fix-security-bug/run2_druid-patching-checklist Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 jackson empty key vulnerability
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/fix-security-bug/run2_jackson-empty-key-vulnerability Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/github-repo-analytics/run2_gh_api_curl_master Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/github-repo-analytics/run2_robust_community_metrics Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.