agentsclimarketplace

Evals & benchmarks

944 rows, by stacks then stars

508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid

  • Aiconfig online evals

    launchdarkly/ai-tooling/skills/agentcontrol/aiconfig-online-evals Skill

    no license25 repo

    LaunchDarkly's official AI tooling

  • Online evals

    launchdarkly/ai-tooling/skills/agentcontrol/online-evals Skill

    no license25 repo

    LaunchDarkly's official AI tooling

  • Gstack benchmark models

    charlieviettq/awesome-agent-skill/.claude/skills/gstack-benchmark-models Skill

    24 repo

    Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

  • Gstack benchmark

    charlieviettq/awesome-agent-skill/.claude/skills/gstack-benchmark Skill

    24 repo

    Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

  • Agent evaluation

    charlieviettq/awesome-agent-skill/.cursor/skills/ai-agent-systems/agent-evaluation Skill

    24 repo

    Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

  • Benchmark models

    charlieviettq/awesome-agent-skill/.cursor/skills/gstack/browser-qa/benchmark-models Skill

    24 repo

    Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

  • Benchmark

    charlieviettq/awesome-agent-skill/.cursor/skills/gstack/browser-qa/benchmark Skill

    24 repo

    Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

  • Benchmark paper

    ShaishavMaisuria/research-paper-lifecycle-skills/skills/benchmark-paper Skill

    23 repo

    42 AI agent skills for literature review, academic writing, citation verification, conference submission, rebuttal, publication, and presentations.

  • Eval grader

    byerlikaya/claude-starter-kit/claude-starter/skills/eval-grader Skill

    22 repo

    Enterprise engineering workflow for Claude Code — not just prompts. AI agents that plan, build, audit, and ship with security gates, privacy checks, and approval-controlled commits. Safely adopt it into new or existing repositories.

  • Eval grader

    byerlikaya/claude-starter-kit/plugin/skills/eval-grader Skill

    22 repo

    Enterprise engineering workflow for Claude Code — not just prompts. AI agents that plan, build, audit, and ship with security gates, privacy checks, and approval-controlled commits. Safely adopt it into new or existing repositories.

  • Skill creator

    ur-grue/autopunk-media-skills/.claude/skills/skill-creator Skill

    21 repo

    394 free AI skills and 6 agents for media professionals — journalists, producers, podcasters, YouTubers. Quality-tested. Works with Claude, ChatGPT, Cursor, Codex CLI, and more.

  • Fluentcrm funnel benchmark

    Lonsdale201/wp-agent-skills/fluentcrm/fluentcrm-funnel-benchmark Skill

    21 repo

    A community-maintained collection of agent skills for WordPress plugin and theme development.

  • Llm overreliance hallucination

    ShulkwiSEC/bb-huge/skills/curated/llm-overreliance-hallucination Skill

    21 repo

    bb-huge 🤗 , Personal bug bounty findings hub and bug bounty orchestration for multiple agents

  • Ds eval

    Dataslayer-AI/Marketing-skills/skills/ds-eval Skill

    19 repo

    Marketing agent skills powered by real data. Connect Claude Code to your actual Google Ads, GA4, Search Console, Meta Ads, LinkedIn Ads and 50+ platforms via Dataslayer MCP — no copy-pasting required.

  • Skill creator

    archubbuck/workspace-architect/assets/skills/skill-creator Skill

    18 repo

    Workspace Architect is a zero-friction CLI tool that provides curated collections of specialized agents, instructions, and prompts to supercharge your GitHub Copilot experience.

  • Hallucination red team

    zgbrenner/agentcounsel/skills/legal-methodology/hallucination-red-team Skill

    17 repo

    Open-source, AI-agnostic skills for legal teams.

  • Recipe eval prompt

    shinpr/rashomon/skills/recipe-eval-prompt Skill

    17 repo

    Measure prompt and skill improvements with blind A/B comparison.

  • Recipe eval skill

    shinpr/rashomon/skills/recipe-eval-skill Skill

    17 repo

    Measure prompt and skill improvements with blind A/B comparison.

  • Skill creator

    codeclawd/fable-mode/skills/skill-creator Skill

    no license16 repo

    Run Claude Fable 5 on Opus 4.8 in Claude Code. The Mythos-class model pulled by export controls — brought back as a native agentic distillation (FABLE_CODE.md), measured playbook, verification hooks, and design/MCP/test skills. "Fable 5 Lite," done right.

  • Borderline eval usage 15

    Teycir/SkillsGuard/testskills/realworld/skills/borderline-eval-usage-15 Skill

    no license15 repo

    Static security scanner for AI agent skill packages. Detects malicious SKILL.md files and bundled scripts before they run.

  • Borderline eval usage 21

    Teycir/SkillsGuard/testskills/realworld/skills/borderline-eval-usage-21 Skill

    no license15 repo

    Static security scanner for AI agent skill packages. Detects malicious SKILL.md files and bundled scripts before they run.

  • Borderline eval usage 3

    Teycir/SkillsGuard/testskills/realworld/skills/borderline-eval-usage-3 Skill

    no license15 repo

    Static security scanner for AI agent skill packages. Detects malicious SKILL.md files and bundled scripts before they run.

  • Borderline eval usage 9

    Teycir/SkillsGuard/testskills/realworld/skills/borderline-eval-usage-9 Skill

    no license15 repo

    Static security scanner for AI agent skill packages. Detects malicious SKILL.md files and bundled scripts before they run.

  • Iblai api agent eval

    iblai/api/skills/iblai-api-agent-eval Skill

    15 repo

    Agent skills + a chat MCP server to operate the ibl.ai platform via its REST API. Install: npx skills add iblai/api

  • Agentv dev

    EntityProcess/agentv/plugins/agentv-dev/skills/agentv-dev Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Agentv bench

    EntityProcess/agentv/skills-data/agentv-bench Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Agentv eval migrations

    EntityProcess/agentv/skills-data/agentv-eval-migrations Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Agentv eval review

    EntityProcess/agentv/skills-data/agentv-eval-review Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Agentv eval writer

    EntityProcess/agentv/skills-data/agentv-eval-writer Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Agentv governance

    EntityProcess/agentv/skills-data/agentv-governance Skill

    15 repo

    Light-weight AI agent evaluation and optimization framework

  • Benchmark table parsing and aggregation

    HolobiomicsLab/asb-skill-collections/collections/epigenomics/v1/skills/benchmark-table-parsing-and-aggregation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Epic array simulation benchmark

    HolobiomicsLab/asb-skill-collections/collections/epigenomics/v1/skills/epic-array-simulation-benchmark Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Cross validation benchmark evaluation

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v1/skills/cross-validation-benchmark-evaluation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark analysis

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v1/skills/mass-spectrometry-benchmark-analysis Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark design

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v1/skills/mass-spectrometry-benchmark-design Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Annotation benchmark performance evaluation

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/annotation-benchmark-performance-evaluation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark dataset generation from reference metabolites

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/benchmark-dataset-generation-from-reference-metabolites Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark harness execution

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/benchmark-harness-execution Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark table generation and reporting

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/benchmark-table-generation-and-reporting Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Cross validation benchmark evaluation

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/cross-validation-benchmark-evaluation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Gallery benchmark data extraction

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/gallery-benchmark-data-extraction Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectral library benchmark execution

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/mass-spectral-library-benchmark-execution Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark analysis

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/mass-spectrometry-benchmark-analysis Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark design

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/mass-spectrometry-benchmark-design Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Metabolite benchmark dataset validation

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/metabolite-benchmark-dataset-validation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Metabolite benchmark peak matching

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/metabolite-benchmark-peak-matching Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Molecular model benchmark comparison

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/molecular-model-benchmark-comparison Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Performance benchmark execution

    HolobiomicsLab/asb-skill-collections/collections/metabolomics/v2/skills/performance-benchmark-execution Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Transcript quantification benchmark comparison

    HolobiomicsLab/asb-skill-collections/collections/transcriptomics/v1/skills/transcript-quantification-benchmark-comparison Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Annotation benchmark performance evaluation

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/lc-ms/skills/annotation-benchmark-performance-evaluation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark harness execution

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/lc-ms/skills/benchmark-harness-execution Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark analysis

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/lc-ms/skills/mass-spectrometry-benchmark-analysis Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark dataset generation from reference metabolites

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/benchmark-dataset-generation-from-reference-metabolites Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark table generation and reporting

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/benchmark-table-generation-and-reporting Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Gallery benchmark data extraction

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/gallery-benchmark-data-extraction Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectral library benchmark execution

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/mass-spectral-library-benchmark-execution Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Mass spectrometry benchmark design

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/mass-spectrometry-benchmark-design Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Metabolite benchmark dataset validation

    HolobiomicsLab/asb-skill-collections/packs/metabolomics/ms-generic/skills/metabolite-benchmark-dataset-validation Skill

    15 repo

    Curated, evidence-grounded skill and software-tool collections for scientific AI agents, generated by the AgenticScienceBuilder

  • Benchmark

    ngocsangyem/MeowKit/.claude/skills/benchmark Skill

    14 repo

    Production ready. AI Agent Workflow System for Claude Code

  • Mk benchmark

    ngocsangyem/MeowKit/packages/mewkit/src/migrate/modules/codex/root/.agents/skills/mk-benchmark Skill

    14 repo

    Production ready. AI Agent Workflow System for Claude Code