agentsclimarketplace

Evals & benchmarks

944 rows, by stacks then stars

508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid

  • Benchmark judge run

    LoogacyStudio/skills/.github/skills/benchmark-judge-run Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark review candidate

    LoogacyStudio/skills/.github/skills/benchmark-review-candidate Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Skill creator

    VRIL-LABS/skill-jam/skills/ai-ml/claude-plugins-official-main/claude-plugins-official-main/plugins/skill-creator/skills/skill-creator Skill

    no license0 repo

    Welcome to the skill-jam ☄️🏀

  • Skill creator

    Anoxxx/skillcache/examples/skill-creator Skill

    no license0 repo

    MRU attention protocol for skills that learn from every use — zero compute

  • Benchmark email automation

    adryanmoldokkr32-pixel/awesome-claude-skills/composio-skills/benchmark-email-automation Skill

    no license0 repo

    A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows

  • Skill quality

    adhi-jp/agent-skills/skills/skill-quality Skill

    0 repo

    Agent skills and eval prompts for vibe-coding plans, review loops, commit messages, prose, and Minecraft modding.

  • Cis benchmark audit

    vahagn-madatyan/netsec-skills-suite/skills/cis-benchmark-audit Skill

    0 repo

    The Complete Network & Network Security Skills Suite

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/data-scientist/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/ml/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Ai evaluation

    SWEStash/swe-workflow-skills/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Skill creator

    ThiagoPanini/aidriven/.claude/skills/skill-creator Skill

    0 repo

    🔩 A Python CLI tool for getting and installing AI development resources into software projects.

  • Skill creator

    ThiagoPanini/aidriven/.github/skills/skill-creator Skill

    0 repo

    🔩 A Python CLI tool for getting and installing AI development resources into software projects.

  • Ai eval regression tester

    sisodiabhumca/agent-skills/skills/ai-eval-regression-tester Skill

    0 repo

    Production-Ready Agent Skills : product analytics, growth experiments, CRM, research synthesis, postmortems, data contracts, SaaS spend, compliance, architecture maps, and LLM eval and many more.

  • Skill creator

    Waterinsoluble-orderplatyctenea746/book2skills/skills/skill-creator Skill

    0 repo

    Transform classic books into AI agent skills for practical, source-based decision making

  • Skill creator

    ploteddie-bit/skills/skill-creator Skill

    no license0 repo

    Collection de compétences agentiques modulables pour Kimi, Aegis et assistants IA

  • Agent platform eval flywheel

    h3y6e/agent-skills/.vendor/skills/agent-platform-eval-flywheel Skill

    0 repo

    my agent skills

  • System benchmark

    h3y6e/agent-skills/.vendor/skills/system-benchmark Skill

    0 repo

    my agent skills

  • Waxa eval

    h3y6e/agent-skills/.vendor/skills/waxa-eval Skill

    0 repo

    my agent skills

  • Evals

    mj-deving/pai-skills/skills/Utilities/Evals Skill

    no license0 repo

    Curated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).

  • Utilities

    mj-deving/pai-skills/skills/Utilities Skill

    no license0 repo

    Curated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).

  • Benchmark

    kjuhwa/skills-hub/skills/cli/gstack/benchmark Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Llm as judge loop

    kjuhwa/skills-hub/skills/llm-agents/llm-as-judge-loop Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Tts trailing silence and hallucination trim

    kjuhwa/skills-hub/skills/ml-ops/tts-trailing-silence-and-hallucination-trim Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Threadpool benchmark with warmup

    kjuhwa/skills-hub/skills/observability/threadpool-benchmark-with-warmup Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Benchmark

    kjuhwa/skills-hub/skills/testing/benchmark Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Lm evaluation harness

    fengluisobel/ask-dongfeng-hermes/skills/mlops/evaluation/lm-evaluation-harness Skill

    0 repo

    Hermes Agent distribution with Ask DongFeng bundled as a built-in control-loop framework skill

  • Skill creator

    Levix0501/next-agent-rails/.agents/skills/skill-creator Skill

    0 repo

    Opinionated Next.js 16 starter for building with AI agents — ship features fast while architecture and directory structure stay under control: fixed structure, enforced conventions, curated libraries, preinstalled skills, native toolchain.

  • Eval runner

    MarchRory/liushi_agent_settings/.agents/skills/eval-runner Skill

    no license0 repo

    Portable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.

  • Eval runner

    MarchRory/liushi_agent_settings/packages/harness/src/.agents/skills/eval-runner Skill

    no license0 repo

    Portable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.

  • Gsd eval review

    bitrails-dev/skills/gsd/skills/gsd-eval-review Skill

    no license0 repo

    Handy AI skills

  • Skill eval pipeline

    pngdeity/apm-user-repository/packages/skill-eval-pipeline/.apm/skills/skill-eval-pipeline Skill

    0 repo

    Personal APM marketplace — skills, prompts, agents, and instructions for AI coding agents

  • Skill creator

    marktantongco/opencodelinux/skills/skill-creator Skill

    no license0 repo

    Merged opencode config + agents + accomplishments showcase — 217 skills, 17 agent profiles, 6 MCP servers, 78-server registry

  • Skill creator

    cmdecker95/skills/skills/skill-creator Skill

    no license0 repo

    🦾 My own little curation of skills

  • Eval best practices

    justinramos101/agent-skill-kit/skills/eval-best-practices Skill

    0 repo

    Battle-tested Agent Skills for coding agents — source-grounded, failure-driven heuristics that audit and design real surfaces. Install with npx skills.

  • Eval driven development

    jacob-balslev/skills/skills/ai-engineering/eval-driven-development Skill

    0 repo

    Public Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills

  • Prompt injection defense

    jacob-balslev/skills/skills/ai-engineering/prompt-injection-defense Skill

    0 repo

    Public Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills

  • Skill creator

    MauroProto/mis-skills/external/anthropics-skills/skill-creator Skill

    0 repo

    Skills personales de diseño frontend (taste, brutalist, minimalist, soft, etc.) + cross-agent sync para Claude Code, Codex, Cursor, Continue, Antigravity. SKILL.md compatible.

  • Negative skill eval writer

    vibesec-advisory/vibesec-advisory-skill-library/SKILLS/ai-workflow-eval-and-lifecycle-review/skills/negative-skill-eval-writer Skill

    0 repo

    Free AI workflow skill libraries for GTM teams, with implementation patterns, guardrails, and evals.

  • Eval case converter

    vibesec-advisory/vibesec-advisory-skill-library/SKILLS/workflow-failure-label-review/skills/eval-case-converter Skill

    0 repo

    Free AI workflow skill libraries for GTM teams, with implementation patterns, guardrails, and evals.

  • Traigent eval audit

    Traigent/traigent-skills/skills/traigent-eval-audit Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Traigent eval build

    Traigent/traigent-skills/skills/traigent-eval-build Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Traigent eval choose metric

    Traigent/traigent-skills/skills/traigent-eval-choose-metric Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Llm observability and evals

    wanghong5233/agent-engineering-kit/cursor/.cursor/skills/llm-observability-and-evals Skill

    0 repo

    A reusable, production-grade .cursor/ engineering package for Cursor / Claude Code / Agent IDEs. Rules, skills, commands, and deterministic safety hooks extracted from a real Agent project.

  • Skill creator

    jae-labs/skills/skills/skill-creator Skill

    0 repo

    Agent Skills.

  • Evaluation lm evaluation harness

    percymcn/agent-cookbook/skills/mlops/evaluation-lm-evaluation-harness Skill

    0 repo

    Production-ready AI agent skills, playbooks, and workflows from the pharma6 automation lab

  • Eval loops

    K-9Nine/skipthespec/eval-loops Skill

    0 repo

    Claude Code skills for prototype-first, autonomous AI development — skip the PRD, dogfood the prototype, run automated eval loops, ship behind flags. Anthropic's playbook as installable Agent Skills.

  • Inference model eval

    cfregly/gpu-perf-tune/plugins/profile-and-optimize/skills/inference-model-eval Skill

    0 repo

    31 GPU inference profiling and optimization skills for Claude Code, with a bundled MCP server

  • Skill benchmark

    TheophilusChinomona/idev/skills/skill-benchmark Skill

    0 repo

    Claude Code plugin: token-optimized dev workflow — cached project context, pattern scanners, build & wiring verification, session persistence, self-review. 23 skills, 3 agents, 4 commands.

  • Agent eval

    goharabbas321/zeoel/skills/agent-eval Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Benchmark

    goharabbas321/zeoel/skills/benchmark Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Eval harness

    goharabbas321/zeoel/skills/eval-harness Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Authoring skills with evals

    vemodalen-x/VEMO_SKILLS/skills/governance/authoring-skills-with-evals Skill

    0 repo

    Public reusable skill hub for agent workflows

  • Rendering html eval reports

    vemodalen-x/VEMO_SKILLS/skills/orchestration/rendering-html-eval-reports Skill

    0 repo

    Public reusable skill hub for agent workflows

  • Skill evaluator

    takehiro177/skill-evaluator/skills/skill-evaluator Skill

    0 repo

    Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

  • Drift eval

    Yco-0314/strata/skills/l5-meta/drift-eval Skill

    0 repo

    Strata — a self-measuring skills framework for Claude Code & Codex: best-of-breed engineering skills stacked by altitude, where every skill has to prove it moves a measured number.

  • Skill eval

    Yco-0314/strata/skills/l5-meta/skill-eval Skill

    0 repo

    Strata — a self-measuring skills framework for Claude Code & Codex: best-of-breed engineering skills stacked by altitude, where every skill has to prove it moves a measured number.

  • Promptfoo evals

    CohenD/cc-stack/.claude/skills/promptfoo-evals Skill

    no license0 repo

    Pre-wired Claude Code workspace: Next.js 16 + AI SDK 7 + shadcn/ui + Tailwind v4 + DuckDB, with Skills & MCP servers ready on clone.

  • Deid reid harness

    Hefrock/agent-skills/skills/deid-reid-harness Skill

    0 repo

    Portable AI agent skills built on the open SKILL.md standard — self-contained capabilities any compatible agent can discover and load on demand, usable across Claude, Codex, Gemini CLI, Cursor, and GitHub Copilot.

  • Anti hallucination

    udsy19/.claude/skills/anti-hallucination Skill

    no license0 repo

    A production-grade .claude folder for Claude Code: 40 lifecycle skills, agents, commands, and hooks — with reflexive skill routing and always-on engineering disciplines (verify, no-bloat, continuous-git, memory, ship-fast).

  • Cross eval

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Business & Ops/cross-eval Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.