agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Incident to eval

    hamidettefagh/incident-to-eval Skill

    0

    A Claude skill that turns a production AI agent incident into a golden eval case. The loop that feeds the two gates.

  • Incident to eval

    hamidettefagh/two-gates/skills/incident-to-eval Skill

    0 repo

    A Claude Code plugin that installs the two-gate method for production AI agents: a design gate, a ship gate, and an incident-to-eval loop that turns failures into regression tests.

  • Llm eval harness and scoring pipeline

    HamzaYM/reliable-ai-skills/skills/evals-and-scoring/llm-eval-harness-and-scoring-pipeline Skill

    29 days old0 repo

    A curated library of production-tested skills for building reliable AI systems with Claude Code. By Hamza Malik, hamz.ai.

  • Agent eval

    Hefrock/agent-skills/skills/agent-eval Skill

    0 repo

    Portable AI agent skills built on the open SKILL.md standard — self-contained capabilities any compatible agent can discover and load on demand, usable across Claude, Codex, Gemini CLI, Cursor, and GitHub Copilot.

  • Deid reid harness

    Hefrock/agent-skills/skills/deid-reid-harness Skill

    0 repo

    Portable AI agent skills built on the open SKILL.md standard — self-contained capabilities any compatible agent can discover and load on demand, usable across Claude, Codex, Gemini CLI, Cursor, and GitHub Copilot.

  • Hermes swarm benchmark

    hzang12345-ship-it/hermes-swarm-benchmark Skill

    0

    Concurrent-agent benchmark suite packaged as a Hermes skill. Markdown REPORT.md output

  • Senior eval engineer

    iamdemetris/lude-kit/skills/personas/senior-eval-engineer Skill

    no license0 repo

    A curated library of senior grade Agent Skills and subagents for Claude Code and OpenAI Codex. 70 skills, 30 dispatchable subagents, designed for multi agent orchestration.

  • Blender agent benchmark

    ifBars/blender-agent-studio/plugins/blender-agent-studio/skills/blender-agent-benchmark Skill

    12 days old0 repo

    Codex plugin for reproducible Blender modeling, validation, animation, MCP tooling, and agent benchmarking

  • Eval driven development

    jacob-balslev/skills/skills/ai-engineering/eval-driven-development Skill

    0 repo

    Public Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills

  • Prompt injection defense

    jacob-balslev/skills/skills/ai-engineering/prompt-injection-defense Skill

    0 repo

    Public Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills

  • Skill creator

    jae-labs/skills/skills/skill-creator Skill

    0 repo

    Agent Skills.

  • Cdfi peer benchmark

    Jaypatel1511/cdfi-superpowers/skills/cdfi-peer-benchmark Skill

    28 days old0 repo

    AI skills for NMTC eligibility, bank-CDFI peer benchmarking & HMDA analysis — grounded in audited PyPI tools, not hallucinated.

  • Eval harness

    Jorgut/context-quality-suite/skills/eval-harness Skill

    14 days old0 repo
  • Benchmark

    joshuaporth/appsec-skill/.cursor/skills/benchmark Skill

    0 repo

    Portable secure code review skill for AI coding agents — OWASP/CWE coverage, structured findings, and remediation guidance. Works with Cursor, Claude Code, Kiro, and Open Agent Skills.

  • Product ai evals

    jpoindexter/product-management-skills/skills/product-ai-evals Skill

    14 days old0 repo

    Operational product-management skills and a /pm router for Codex and Claude.

  • Contrastive authority eval

    jpoindexter/reward-seeking-safety-skills/skills/contrastive-authority-eval Skill

    15 days old0 repo

    Agent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.

  • Eval awareness red team

    jpoindexter/reward-seeking-safety-skills/skills/eval-awareness-red-team Skill

    15 days old0 repo

    Agent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.

  • Eval best practices

    justinramos101/agent-skill-kit/skills/eval-best-practices Skill

    0 repo

    Battle-tested Agent Skills for coding agents — source-grounded, failure-driven heuristics that audit and design real surfaces. Install with npx skills.

  • Eval loops

    K-9Nine/skipthespec/eval-loops Skill

    0 repo

    Claude Code skills for prototype-first, autonomous AI development — skip the PRD, dogfood the prototype, run automated eval loops, ship behind flags. Anthropic's playbook as installable Agent Skills.

  • Native agent evals

    Kbediako/evergreen-codex-skills/skills/native-agent-evals Skill

    21 days old0 repo

    Evergreen, generally useful Codex skills designed to stay valuable as models improve.

  • Skill creator

    kennyolofsson23-netizen/claude-code-config/skills/skill-creator Skill

    no license0 repo

    The most comprehensive Claude Code setup — 29 skills, 9 plugins, 13 hooks, 8 agents, 8 commands. Self-documenting, self-improving, CLI-first. Includes self-interview prompt for personalization.

  • Benchmark

    kjuhwa/skills-hub/skills/cli/gstack/benchmark Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Llm as judge loop

    kjuhwa/skills-hub/skills/llm-agents/llm-as-judge-loop Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Tts trailing silence and hallucination trim

    kjuhwa/skills-hub/skills/ml-ops/tts-trailing-silence-and-hallucination-trim Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Threadpool benchmark with warmup

    kjuhwa/skills-hub/skills/observability/threadpool-benchmark-with-warmup Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Benchmark

    kjuhwa/skills-hub/skills/testing/benchmark Skill

    0 repo

    Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

  • Case 00228

    knownasnaffy/prompthound/dataset/case_00228 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 00490

    knownasnaffy/prompthound/dataset/case_00490 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 01169

    knownasnaffy/prompthound/dataset/case_01169 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 01745

    knownasnaffy/prompthound/dataset/case_01745 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 02998

    knownasnaffy/prompthound/dataset/case_02998 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 03442

    knownasnaffy/prompthound/dataset/case_03442 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 03796

    knownasnaffy/prompthound/dataset/case_03796 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Case 05375

    knownasnaffy/prompthound/dataset/case_05375 Skill

    no license0 repo

    A fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.

  • Competitor benchmark report

    kochellenk-afk/google-ads-diagnostic-toolkit/skills/competitor-benchmark-report Skill

    19 days old0 repo

    10 production Claude Skills covering the full Google Ads diagnostic lifecycle: waste, spikes, Quality Score, budgets, reporting

  • Skill creator

    lawwu/skills-marketplace/{{cookiecutter.repo_name}}/plugins/{{cookiecutter.plugin_name}}/skills/skill-creator Skill

    no license0 repo

    Cookiecutter Template for an Agent Skills Marketplace

  • Ai eval

    lazymac2x/ai-eval-api/io.github.lazymac2x/ai-eval MCP server

    no license0 repo

    Cloudflare Workers MCP server: ai-eval

  • Skill creator

    learn-with-santosh/claude-master-skills/skills/skill-creator Skill

    no license0 repo

    A curated collection of specialized skills and workflows designed to enhance Claude's capabilities in specific domains. These skills provide frameworks, psychological triggers, and structured processes to deliver high-quality, professional results.

  • Skill creator

    Levix0501/next-agent-rails/.agents/skills/skill-creator Skill

    0 repo

    Opinionated Next.js 16 starter for building with AI agents — ship features fast while architecture and directory structure stay under control: fixed structure, enforced conventions, curated libraries, preinstalled skills, native toolchain.

  • Arbor

    lhbsaa/apex-discovery/skills/arbor Skill

    15 days old0 repo

    Scientific research extension for apex-unified - 149 scientific skills for biology, chemistry, medicine, materials research. Powered by Scientific Agent Skills (31.4k star).

  • Autoresearch

    libenxier-beep/codex-custom-skills/skills/autoresearch Skill

    0 repo

    Production-grade Codex skills with explicit triggers, deterministic validation, and reusable AI agent workflows.

  • Test runner

    Liber1917/agent-skill-infra/skills/test-runner Skill

    no license0 repo

    Agent Skill Infrastructure: quality check, behavior test runner, version awareness

  • Benchmark author item

    LoogacyStudio/skills/.github/skills/benchmark-author-item Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark author rerun manifest

    LoogacyStudio/skills/.github/skills/benchmark-author-rerun-manifest Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark author skill evals

    LoogacyStudio/skills/.github/skills/benchmark-author-skill-evals Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark author variants

    LoogacyStudio/skills/.github/skills/benchmark-author-variants Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark cluster failures

    LoogacyStudio/skills/.github/skills/benchmark-cluster-failures Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark core

    LoogacyStudio/skills/.github/skills/benchmark-core Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark judge run

    LoogacyStudio/skills/.github/skills/benchmark-judge-run Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Benchmark review candidate

    LoogacyStudio/skills/.github/skills/benchmark-review-candidate Skill

    0 repo

    This repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.

  • Eval runner

    MarchRory/liushi_agent_settings/.agents/skills/eval-runner Skill

    no license0 repo

    Portable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.

  • Eval runner

    MarchRory/liushi_agent_settings/packages/harness/src/.agents/skills/eval-runner Skill

    no license0 repo

    Portable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.

  • Skill creator

    marktantongco/opencodelinux/skills/skill-creator Skill

    no license0 repo

    Merged opencode config + agents + accomplishments showcase — 217 skills, 17 agent profiles, 6 MCP servers, 78-server registry

  • Output scoring en

    matematicsolutions/awesome-matematic-skills-en/verification-foundation/skills/output-scoring-en Skill

    no license0 repo

    English hub of method-neutral legal AI skills (verification core, content quality, EU law) - bundle model. Polish-jurisdiction skills live in awesome-matematic-skills-pl.

  • Karvey benchmark models

    MauricioQuezadaHaintech/karvey/plugins/karvey/skills/karvey-benchmark-models Skill

    0 repo

    Karvey — método spec-driven development agnóstico de stack (Afán, selknam). Plugin de Claude Code. © HainTech, Apache 2.0.

  • Skill creator

    MauroProto/mis-skills/external/anthropics-skills/skill-creator Skill

    0 repo

    Skills personales de diseño frontend (taste, brutalist, minimalist, soft, etc.) + cross-agent sync para Claude Code, Codex, Cursor, Continue, Antigravity. SKILL.md compatible.

  • Thesis theory evaluator

    mcmespinaa/thesis-theory-evaluator Skill

    0

    Claude Code skill that evaluates a thesis Theoretical Framework chapter against the OL646E (Malmö University SALSU) course rubric. Companion to thesis-anchoring-map.

  • Agent eval harness

    megandmartin/agent-skills-repo/skills/agent-mastery/agent-eval-harness Skill

    12 days old0 repo

    75 production-grade agent skills for Hermes Agent + Paperclip — research, write, organize, earn, and run an AI workforce. Every skill passes a QA gate with hard safety rails. Built by Gen AI Hub.

  • Agent eval

    Methasit-Pun/ts-ddd-clean-architecture/.claude/skills/agent-eval Skill

    0 repo

    TypeScript DDD Clean Architecture Skills

  • Skill creator

    michelve/hugin-v0/skills/skill-creator Skill

    0 repo

    A Claude Code plugin packaging 23 skills, 8 agents, 5 event hooks, and 7 MCP servers for full-stack development with React 19, TypeScript, Express, Prisma, Tailwind CSS v4, and shadcn/ui.