The catalog
325,949 rows, by stacks then stars
Run2 python docx nested table recursion
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/offer-letter-generator/run2_python-docx-nested-table-recursion Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 pdf text extraction for classification
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/organize-messy-files/run3_pdf-text-extraction-for-classification Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 read python tokenizer source
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/python-scala-translation/run3_read-python-tokenizer-source Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/python-scala-translation/run3_read-scala-test-spec Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 scala tokenizer implementation guide
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/python-scala-translation/run3_scala-tokenizer-implementation-guide Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 email json parsing and meeting extraction
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/schedule-planning/run1_email-json-parsing-and-meeting-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/schedule-planning/run1_pdf-calendar-parsing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/stock-data-visualization/run3_skill-1 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 glm observation merge rmse
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/temperature-simulation/run3_glm-observation-merge-rmse Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/temperature-simulation/run3_glm-output-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 database search ohio trip
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/travel-planning/run3_database-search-ohio-trip Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 ohio trip itinerary planning
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/travel-planning/run3_ohio-trip-itinerary-planning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 convert keyframes to grayscale
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/video-object-counting/run3_convert-keyframes-to-grayscale Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 count objects in frames using template matching
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/video-object-counting/run3_count-objects-in-frames-using-template-matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/video-object-counting/run3_generate-counting-csv Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 video keyframe extraction
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/video-object-counting/run3_video-keyframe-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 excel lookup formulas gnumeric compatible
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/weighted-gdp-calculation/run3_excel-lookup-formulas-gnumeric-compatible Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 gdp weighted mean net exports gcc
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/weighted-gdp-calculation/run3_gdp-weighted-mean-net-exports-gcc Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-opus-4-6/weighted-gdp-calculation/run3_inspect-excel-structure Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_anthropic-brand-tokens Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/chinese-poem-generator/run3_qilv-tonal-patterns Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 fill sc100 small claims form
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/court-form-filling/run3_fill-sc100-small-claims-form Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/court-form-filling/run3_inspect-pdf-fields Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 dbscan custom metric clustering
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run1_dbscan-custom-metric-clustering Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run1_greedy-centroid-matching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 pareto frontier computation
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run1_pareto-frontier-computation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_find-trivy-and-cache Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_run-trivy-audit Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 find earthquake farthest from pacific boundary
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run3_find-earthquake-farthest-from-pacific-boundary Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run3_load-and-inspect-geodata Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/enterprise-information-search/run3_skill-1 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/financial-analysis/run3_count-renaissance-stocks Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/financial-analysis/run3_get-fund-holdings Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/financial-analysis/run3_load-13f-coverpage Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/fix-security-bug/run3_skill-1 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/github-repo-analytics/run3_github-search-issues Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/github-repo-analytics/run3_github-search-prs Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/github-repo-analytics/run3_write-report-json Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 check and setup python310
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run3_check-and-setup-python310 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run3_implement-simpo-loss Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 install simpo dependencies python310
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run3_install-simpo-dependencies-python310 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 run simpo with python310 and log
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/nlp-paper-reproduction/run3_run-simpo-with-python310-and-log Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/offer-letter-generator/run3_fill-docx-template Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 organize files by subject
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/organize-messy-files/run3_organize-files-by-subject Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 python to scala translation patterns
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/python-scala-translation/run1_python-to-scala-translation-patterns Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 scala tokenizer domain model
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/python-scala-translation/run1_scala-tokenizer-domain-model Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/schedule-planning/run1_json-email-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/schedule-planning/run1_pdf-calendar-parsing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 build d3 bubble chart with force simulation
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_build-d3-bubble-chart-with-force-simulation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 build interactive data table with bubble sync
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_build-interactive-data-table-with-bubble-sync Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 copy files and create directories
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_copy-files-and-create-directories Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_download-d3-v6 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_list-directory-files Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_read-csv-files Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 write single page web app files
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/stock-data-visualization/run3_write-single-page-web-app-files Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/temperature-simulation/run2_glm_lake_mendota_setup Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/temperature-simulation/run2_glm_output_parsing Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/travel-planning/run3_itinerary-formatting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/travel-planning/run3_itinerary-planning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-sonnet-4-6/travel-planning/run3_itinerary-validation Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.