agentsclimarketplace

The catalog

325,949 rows, by stacks then stars

  • Run2 csv reporting

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_csv-reporting Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 cvss score

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_cvss-score Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 trivy offline

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/dependency-vulnerability-check/run2_trivy-offline Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 antimeridian handling

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_antimeridian-handling Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 geopandas geospatial distance

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_geopandas-geospatial-distance Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 plate boundary analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/earthquake-plate-calculation/run2_plate-boundary-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 enterprise data retrieval

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/enterprise-information-search/run2_enterprise-data-retrieval Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 slack analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/enterprise-information-search/run2_slack-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 13f data analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/financial-analysis/run2_13f-data-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 druid cve 2021 25646

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 jackson inject security

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/fix-security-bug/run2_jackson-inject-security Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 data analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/github-repo-analytics/run2_data-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 gh rest api

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/github-repo-analytics/run2_gh-rest-api Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 pytorch preference optimization

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/nlp-paper-reproduction/run2_pytorch-preference-optimization Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 simpo loss

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/nlp-paper-reproduction/run2_simpo-loss Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 python docx split placeholders

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/offer-letter-generator/run2_python-docx-split-placeholders Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 document text extraction

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/organize-messy-files/run2_document-text-extraction Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 keyword classification

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/organize-messy-files/run2_keyword-classification Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 scala circe json

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-circe-json Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 scala idiomatic patterns

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-idiomatic-patterns Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 scala immutable builder

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-immutable-builder Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 scala option handling

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-option-handling Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 scala variance

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/python-scala-translation/run2_scala-variance Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 meeting scheduling

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_meeting-scheduling Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 pdf calendar parsing

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_pdf-calendar-parsing Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 timezone conversion

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/schedule-planning/run2_timezone-conversion Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 d3 force bubble

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/stock-data-visualization/run2_d3-force-bubble Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 d3 interactive table

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/stock-data-visualization/run2_d3-interactive-table Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 glm calibration strategy

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_glm-calibration-strategy Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 glm lake modeling

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_glm-lake-modeling Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 netcdf obs matching

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/temperature-simulation/run2_netcdf-obs-matching Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 budget validation

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_budget-validation Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 itinerary planning

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_itinerary-planning Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 travel data extraction

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/travel-planning/run2_travel-data-extraction Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 ffmpeg keyframes

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_ffmpeg-keyframes Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 opencv grayscale

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_opencv-grayscale Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 template matching tuning

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/video-object-counting/run2_template-matching-tuning Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 excel index match

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/weighted-gdp-calculation/run2_excel-index-match Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 excel weighted mean

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-opus-4-6/weighted-gdp-calculation/run2_excel-weighted-mean Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 anthropic brand tokens

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_anthropic-brand-tokens Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 exploded view drawing

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_exploded-view-drawing Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 matplotlib technical poster

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/anthropic-poster-design/run2_matplotlib-technical-poster Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 chinese poetry composition

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/chinese-poem-generator/run2_chinese-poetry-composition Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 qiyan lushi

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/chinese-poem-generator/run2_qiyan-lushi Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 sc100 small claims form

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/court-form-filling/run2_sc100-small-claims-form Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 dbscan clustering

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_dbscan-clustering Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 greedy matching

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_greedy-matching Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 pareto frontier

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dbscan-parameter-tuning/run2_pareto-frontier Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 cvss score extraction

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_cvss-score-extraction Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 trivy offline scanning

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_trivy-offline-scanning Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 vulnerability csv reporting

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/dependency-vulnerability-check/run2_vulnerability-csv-reporting Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 earthquake analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run2_earthquake-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 geopandas plate tectonics

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/earthquake-plate-calculation/run2_geopandas-plate-tectonics Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 competitor analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/enterprise-information-search/run2_competitor-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 enterprise data retrieval

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/enterprise-information-search/run2_enterprise-data-retrieval Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 13f fund analysis

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_13f-fund-analysis Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 fuzzy fund search

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_fuzzy-fund-search Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 holdings comparison

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_holdings-comparison Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 palantir holders

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/financial-analysis/run2_palantir-holders Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

  • Run2 druid cve 2021 25646

    cxcscmu/SkillLearnBench/skills/b2-self-feedback-claude-sonnet-4-6/fix-security-bug/run2_druid-cve-2021-25646 Skill

    75★ repo

    [COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.