agentsclimarketplace

Evals & benchmarks

944 rows, by stacks then stars

508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid

  • Skill creator

    lawwu/skills-marketplace/{{cookiecutter.repo_name}}/plugins/{{cookiecutter.plugin_name}}/skills/skill-creator Skill

    no license0 repo

    Cookiecutter Template for an Agent Skills Marketplace

  • Thesis theory evaluator

    mcmespinaa/thesis-theory-evaluator Skill

    0

    Claude Code skill that evaluates a thesis Theoretical Framework chapter against the OL646E (Malmö University SALSU) course rubric. Companion to thesis-anchoring-map.

  • Agent eval

    Methasit-Pun/ts-ddd-clean-architecture/.claude/skills/agent-eval Skill

    0 repo

    TypeScript DDD Clean Architecture Skills

  • Amortized intelligence

    Sapien3/amortized-intelligence Skill

    0

    A Claude Code skill: spend frontier-model tokens at design time, compile them into artifacts — harnesses, rubrics, evals, skills — that cheap models run at near-zero marginal cost.

  • Epitech pre eval

    epitech-toulouse/epitech-pre-eval Skill

    no license0
  • Skill creator

    wentallout/write-like-a-human/.agents/skills/skill-creator Skill

    no license0 repo

    Messy, opinionated, and aggressively human.

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/ai/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Self eval

    ekreloff/claude-self-eval/skills/self-eval Skill

    no license0 repo

    Honest AI work evaluation for Claude Code — two-axis scoring with anti-inflation mechanisms

  • Eval criteria create

    0SxD/advisor-trackable-eval-skills-v01/skills/eval_criteria_create Skill

    0 repo
  • Benchmark

    joshuaporth/appsec-skill/.cursor/skills/benchmark Skill

    0 repo

    Portable secure code review skill for AI coding agents — OWASP/CWE coverage, structured findings, and remediation guidance. Works with Cursor, Claude Code, Kiro, and Open Agent Skills.

  • Golang benchmark

    guynhsichngeodiec/cc-skills-golang/skills/golang-benchmark Skill

    0 repo

    Extend Go coding assistants with reusable skills for language, testing, security, and observability in production-ready Golang projects

  • Skill creator

    fermeridamagni/skills/.agents/skills/skill-creator Skill

    0 repo

    A collection of AI skills created for me

  • Anti hallucination

    fbsmna-coder/karpathy-pro-max/skills/anti-hallucination Skill

    0 repo

    Stop Claude Code from hallucinating — Karpathy-grade discipline in 8 skills

  • Buyer eval skill

    Zorineinsupportable217/buyer-eval-skill Skill

    0

    Evaluate B2B vendors by talking to AI agents and verifying claims against independent sources to score what is true and what is not

  • Skill ab eval

    cskwork/skill-ab-eval Skill

    0

    Prove whether a SKILL.md actually changes agent behavior (with/without) and which CLI harness does a task best — agent-native subagents or CLI. agentskills.io-compatible. Site: https://cskwork.github.io/skill-ab-eval/

  • Hermes swarm benchmark

    hzang12345-ship-it/hermes-swarm-benchmark Skill

    0

    Concurrent-agent benchmark suite packaged as a Hermes skill. Markdown REPORT.md output

  • Agent eval kit

    claudette-agent/agent-eval-kit Skill

    no license0

    Free mini-eval and scorecard for AI coding agents: Claude Code, Codex, Cursor, Copilot, Windsurf.

  • Clinical note eval

    yashpatil582/agent-skills/clinical-note-eval Skill

    0 repo

    Anthropic Agent Skills packaging shipped eval methods: gaming-resistant clinical-note grounding eval + token-F1 scorer. Stdlib-only, runs offline.

  • Audit the oracle coverage

    XyraSinclair/ideonomy/skills/audit-the-oracle-coverage Skill

    0 repo

    Computational ideonomy: Gunkel's science of ideas as inference-time machinery — a 37-primitive organon, an MDL-ratcheted respiratory engine, and 14 gated agent skills (Claude Code plugin)

  • Autoresearch

    libenxier-beep/codex-custom-skills/skills/autoresearch Skill

    0 repo

    Production-grade Codex skills with explicit triggers, deterministic validation, and reusable AI agent workflows.

  • Sprint eval

    anupam-io/sprint-skills/skills/sprint-eval Skill

    0 repo

    Agent skills that turn a one-line goal into a dependency-ordered sprint of GitHub issues, then run them one at a time — review-each (HITL) or fully unattended (AFK). No cluster, no Linear: just gh, claude, and your repo.

  • Eval harness architect

    satishTheLegend/eval-harness-architect Skill

    0

    Stand up a rigorous, regression-proof evaluation harness for any LLM/agent system from zero — datasets, scorers, CI gates, and drift monitoring.

  • Ai rmf governor

    satishTheLegend/ai-rmf-governor Skill

    0

    Claude Code skill: run NIST AI RMF (Govern/Map/Measure/Manage) on your LLM/ML feature — named failure modes, metric thresholds, an offline eval gate, and signed-off residual risk.

  • Incident to eval

    hamidettefagh/incident-to-eval Skill

    0

    A Claude skill that turns a production AI agent incident into a golden eval case. The loop that feeds the two gates.

  • Add skill

    Paldom/landing-page-builder-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for building high-converting, modern, performant landing pages: design tokens, conversion copywriting, marketing UX, layout patterns, Next.js/React + shadcn implementation, scroll motion, and performance/SEO/accessibility.

  • Add skill

    Paldom/llm-council-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for building and operating LLM councils - multi-model deliberation with anonymized peer review, robust aggregation, calibrated escalation, and defenses against correlated failure.

  • Image benchmark report

    StaryMoon/image-benchmark-report-skill/skills/image-benchmark-report Skill

    0 repo

    Agent Skill for PSNR/SSIM image benchmarks with per-image scores and worst-case reports.

  • Eval authoring

    soltraveler-sri/Steerable-Protocol/skills/eval-authoring Skill

    0 repo

    An open standard for making apps steerable, shipped with a reference runtime and a skill that lets a coding agent implement it in your codebase.

  • Pm trace eval

    Nyota-tree/pm-trace-eval/skills/pm-trace-eval Skill

    0 repo

    PM-facing AI conversation trace eval skill: north-star metrics, 1-5 scores, red/yellow/green, per-turn quote review

  • Add skill

    Paldom/mcp-server-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for designing and implementing Model Context Protocol (MCP) servers - server architecture, tools, resources, prompts, transports, authorization, testing, and deployment.

  • Add skill

    Paldom/screenshooter/.claude/skills/add-skill Skill

    0 repo

    Skills for creating polished screen-recording walkthrough and tutorial videos of web apps, with smooth scripted mouse movements, zoom effects, and explanation cards.

  • Add skill

    Paldom/stealth-browser-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for stealthy, human-like browser automation with Playwright and Camoufox — real-browser agentic browsing that evades bot detection.

  • Add skill

    Paldom/promptimize/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.

  • Prompt rail

    yuanGao0816/prompt-rail Skill

    no license0

    Measured prompt iteration with train/test dual-gate scoring and anti-overfit rails. Agent Skill (SKILL.md).

  • Evalpilot

    avnath13/evalpilot Skill

    0

    Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.

  • Add skill

    Paldom/terminaltor/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for recording, sanitizing, and rendering terminal session demos - capture real command walkthroughs with agents, redact sensitive output, and render polished casts for READMEs and docs.

  • Add skill

    Paldom/noslop/.claude/skills/add-skill Skill

    0 repo

    Agent Skills that strip AI writing tells from copy - em dashes, 'it's not X, it's Y' constructions, hedging, and bloat - so text reads like a human wrote it, with evals to verify.

  • Skill description tuning

    blakebauman/skillist-validator/skills/skill-description-tuning Skill

    0 repo

    Validate AI coding assistant skills against the Agent Skills specification (agentskills.io) — a dependency-free validator with fix hints, plus skills for auditing instruction quality and tuning description triggering.

  • Agent eval harness

    megandmartin/agent-skills-repo/skills/agent-mastery/agent-eval-harness Skill

    0 repo

    75 production-grade agent skills for Hermes Agent + Paperclip — research, write, organize, earn, and run an AI workforce. Every skill passes a QA gate with hard safety rails. Built by Gen AI Hub.

  • Eval harness

    Jorgut/context-quality-suite/skills/eval-harness Skill

    0 repo
  • Puzzletide agent evals

    Caravaca-Labs/puzzletide-cli/skills/puzzletide-agent-evals Skill

    0 repo

    Word search generator, crossword generator, and sudoku generator & solver in one local-first CLI — printable PDF puzzle worksheets, themed word banks, verifiable LLM evals, and agent skills. From the makers of puzzletide.com.

  • Geminiupgradeqa mcp

    clauxel/gemini-upgrade-qa-mcp/com.clauxel.geminiupgradeqa/geminiupgradeqa-mcp MCP server

    0 repo

    Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.

  • Mcp eval runner

    dbsectrainer/mcp-eval-runner/io.github.dbsectrainer/mcp-eval-runner MCP server

    0 repo

    A standardized testing harness for MCP servers and agent workflows

  • Ai eval

    lazymac2x/ai-eval-api/io.github.lazymac2x/ai-eval MCP server

    no license0 repo

    Cloudflare Workers MCP server: ai-eval

  • Skill creator

    CreminiAI/skillpack/templates/builtin-skills/skill-creator Skill

    1,197 repo

    Pack and deploy local AI agents for your team in minutes

  • Xcode build benchmark

    AvdLee/Xcode-Build-Optimization-Agent-Skill/skills/xcode-build-benchmark Skill

    1,195 repo

    An Agent Skill helping you to optimize Xcode incremental and clean builds by running benchmarks and optimizing build settings.

  • Benchmark analyst

    growthack88/growth-marketing-os/skills/benchmark-analyst Skill

    79 repo

    Growth Marketing OS | Mahmoud Omar — open-source AI marketing prompts, Claude skills, agents & growth playbooks (EN + AR)

  • Ai evals

    liqiongyu/lenny_skills_plus/skills/ai-evals Skill

    51 repo

    86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.

  • Research proof

    tonyblu331/research-proof/skills/research-proof Skill

    45 repo

    Pressure-test research claims with falsifiable evidence plans, adversarial checks, frozen verifiers, and proof ledgers.

  • Nasde benchmark creator

    NoesisVision/nasde-toolkit/.claude/skills/nasde-benchmark-creator Skill

    11 repo

    CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.

  • Golang benchmark

    alexastrum/skl/.agents/skills/golang-benchmark Skill

    11 repo

    A lightweight, single-binary Agent Skills CLI manager written in Go

  • Eval surfer

    di37/EvalSurfer/.cursor/skills/eval-surfer Skill

    11 repo

    Skill-first, agent-native evaluation protocol for AI apps

  • Pydantic evals

    Fuenfgeld/pydantic-ai-skills/skills/pydantic-evals Skill

    10 repo

    Production-ready Claude Code skills for building AI agents with Pydantic AI. Includes dependency injection, tools, validators, streaming, multi-agent orchestration, and evaluation framework patterns.

  • Red team eval authoring

    yeaight7/agent-powerups/plugins/agent-evaluation-lab/skills/red-team-eval-authoring Skill

    6 repo

    Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more

  • Ai feature eval harness

    Mozurok/fhorja.dev/.claude/skills/ai-feature-eval-harness Skill

    6 repo

    A workflow operating system for AI-assisted engineering. Task state, decisions, and plans live on disk as files, not in chat history, so context survives across sessions, tools, and restarts.

  • Benchmark models

    timurgaleev/vibestack/skills/benchmark-models Skill

    5 repo

    vibestack is a portable skill pack for AI coding agents. Slash commands like /office-hours, /ship, /investigate, /tdd, /review install once and work across every agent that supports the Agent Skills open standard — Claude Code, Cursor, Kiro, and a growing list of others.

  • Benchmark

    josephneumann/claude-corps/skills/benchmark Skill

    4 repo

    Production workflow framework for Claude Code — multi-agent orchestration, parallel worktrees, autonomous task execution, and compound engineering skills

  • Benchmark e2e

    build-with-dhiraj/ai-workflow-framework-portability-kit/Plugins/vercel-marketplace-source/.claude/skills/benchmark-e2e Skill

    4 repo

    Portable, self-contained snapshot of a complete Claude Code setup — 36 specialist agents, 134 skills, plugins, MCP servers & host tooling. Clone, claude login, run one script, restore the whole orchestration stack in ~20 min.

  • Aicrew benchmark

    AKSoftCode/aicrew/codex-skills/aicrew-benchmark Skill

    3 repo

    TDD-first AI dev pipeline for Claude Code, Cursor, Codex, Gemini CLI & Antigravity. /dev /fix /quick + token-saving hooks. npx aicrew install

  • Floom skill evals

    floomhq/starter/skills/floom-skill-evals Skill

    3 repo

    Floom Starter Pack. 65 hand-picked AI agent skills for Claude Code, Codex, Cursor, Kimi, OpenCode. One install, auto-activates. +16.2pp pass-rate lift.

Catalog - agentscli marketplace