Engineering
99,663 rows, by stacks then stars
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3-flash-preview/fix-security-bug/jackson-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3-flash-preview/fix-security-bug/java-javascript-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3-flash-preview/nlp-paper-reproduction/model-verification-unit-tests Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/anthropic-poster-design/json-configuration-management Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/dbscan-parameter-tuning/pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/dependency-vulnerability-check/trivy-audit Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/financial-analysis/file-data-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/fix-security-bug/apache-druid-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/nlp-paper-reproduction/environment-setup Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/python-scala-translation/scala-data-modeling Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-flash-lite-preview/schedule-planning/calendar-scheduler Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-pro-preview/dbscan-parameter-tuning/pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-pro-preview/dependency-vulnerability-check/trivy-offline-scanning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b1-one-shot-gemini-3.1-pro-preview/fix-security-bug/jackson-druid-cve Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 npm vulnerability scanning enhanced
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/dependency-vulnerability-check/run2_npm-vulnerability-scanning-enhanced Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 druid javascript deserialization security
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_druid-javascript-deserialization-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 jackson security patching
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_jackson-security-patching Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/fix-security-bug/run2_maven-security-build Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/github-repo-analytics/run2_github-rest-api Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 simpo testing and verification
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/nlp-paper-reproduction/run2_simpo-testing-and-verification Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_nml-calibration-workflow Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 parameter optimization strategy
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/temperature-simulation/run2_parameter-optimization-strategy Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-haiku-4-5/travel-planning/run2_itinerary_optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dbscan-parameter-tuning/run2_pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_csv-reporting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_trivy-offline Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_jackson-inject-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/github-repo-analytics/run2_gh-rest-api Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 pytorch preference optimization
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/nlp-paper-reproduction/run2_pytorch-preference-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_travel-data-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_trivy-offline-scanning Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 vulnerability csv reporting
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_vulnerability-csv-reporting Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 meeting scheduler with output
cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/schedule-planning/run2_meeting-scheduler-with-output Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dbscan-parameter-tuning/run2_parallel-orchestration Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dependency-vulnerability-check/run2_trivy-offline Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/dependency-vulnerability-check/run2_vulnerability-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run2 jackson empty key vulnerability
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/fix-security-bug/run2_jackson-empty-key-vulnerability Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3-flash-preview/temperature-simulation/run2_glm-config Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-flash-lite-preview/dbscan-parameter-tuning/run2_pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-flash-lite-preview/dependency-vulnerability-check/run2_trivy_audit Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-flash-lite-preview/nlp-paper-reproduction/run2_pytorch-tensor-ops Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-flash-lite-preview/python-scala-translation/run2_scala-testing-idioms Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-flash-lite-preview/schedule-planning/run2_scheduler Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-pro-preview/dbscan-parameter-tuning/run2_pareto-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-pro-preview/dependency-vulnerability-check/run2_trivy-offline Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-pro-preview/enterprise-information-search/run2_json-data-extraction Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-pro-preview/fix-security-bug/run2_jackson-security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
cxcscmu/SkillLearnBench/skills/b2-self-feedback-gemini-3.1-pro-preview/schedule-planning/run2_scheduler Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run1 pareto frontier optimization
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/dbscan-parameter-tuning/run1_pareto-frontier-optimization Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Deduplicate and Order Vulnerability Records
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/dependency-vulnerability-check/run3_Deduplicate-and-Order-Vulnerability-Records Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Filter Vulnerabilities by Severity Level
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/dependency-vulnerability-check/run3_Filter-Vulnerabilities-by-Severity-Level Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Handle NULL and Missing Values in Vulnerability Data
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/dependency-vulnerability-check/run3_Handle-NULL-and-Missing-Values-in-Vulnerability-Data Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Map Trivy JSON Fields to CSV Columns Accurately
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/dependency-vulnerability-check/run3_Map-Trivy-JSON-Fields-to-CSV-Columns-Accurately Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Filter Holdings by Security Type Stocks Only
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/financial-analysis/run3_Filter-Holdings-by-Security-Type-Stocks-Only Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Search CUSIP for Security
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/financial-analysis/run3_Search-CUSIP-for-Security Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Create and Apply Git Patch Files for Druid Security Fixes
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/fix-security-bug/run3_Create-and-Apply-Git-Patch-Files-for-Druid-Security-Fixes Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Check Repository Dependencies
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/nlp-paper-reproduction/run3_Check-Repository-Dependencies Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.
Run3 Run Unit Test in Python 3 10 Environment
cxcscmu/SkillLearnBench/skills/b3-teacher-feedback-claude-haiku-4-5/nlp-paper-reproduction/run3_Run-Unit-Test-in-Python-3-10-Environment Skill
75★ repo[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.