Evals & benchmarks
944 rows, by stacks then stars
508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid
lawwu/skills-marketplace/{{cookiecutter.repo_name}}/plugins/{{cookiecutter.plugin_name}}/skills/skill-creator Skill
no license0★ repoCookiecutter Template for an Agent Skills Marketplace
mcmespinaa/thesis-theory-evaluator Skill
0★Claude Code skill that evaluates a thesis Theoretical Framework chapter against the OL646E (Malmö University SALSU) course rubric. Companion to thesis-anchoring-map.
Methasit-Pun/ts-ddd-clean-architecture/.claude/skills/agent-eval Skill
0★ repoTypeScript DDD Clean Architecture Skills
Sapien3/amortized-intelligence Skill
0★A Claude Code skill: spend frontier-model tokens at design time, compile them into artifacts — harnesses, rubrics, evals, skills — that cheap models run at near-zero marginal cost.
epitech-toulouse/epitech-pre-eval Skill
no license0★wentallout/write-like-a-human/.agents/skills/skill-creator Skill
no license0★ repoMessy, opinionated, and aggressively human.
SWEStash/swe-workflow-skills/plugins/ai/skills/ai-evaluation Skill
0★ repoA comprehensive SWE workflow, encoded. Might be useful to you too.
ekreloff/claude-self-eval/skills/self-eval Skill
no license0★ repoHonest AI work evaluation for Claude Code — two-axis scoring with anti-inflation mechanisms
0SxD/advisor-trackable-eval-skills-v01/skills/eval_criteria_create Skill
0★ repojoshuaporth/appsec-skill/.cursor/skills/benchmark Skill
0★ repoPortable secure code review skill for AI coding agents — OWASP/CWE coverage, structured findings, and remediation guidance. Works with Cursor, Claude Code, Kiro, and Open Agent Skills.
guynhsichngeodiec/cc-skills-golang/skills/golang-benchmark Skill
0★ repoExtend Go coding assistants with reusable skills for language, testing, security, and observability in production-ready Golang projects
fermeridamagni/skills/.agents/skills/skill-creator Skill
0★ repoA collection of AI skills created for me
fbsmna-coder/karpathy-pro-max/skills/anti-hallucination Skill
0★ repoStop Claude Code from hallucinating — Karpathy-grade discipline in 8 skills
Zorineinsupportable217/buyer-eval-skill Skill
0★Evaluate B2B vendors by talking to AI agents and verifying claims against independent sources to score what is true and what is not
cskwork/skill-ab-eval Skill
0★Prove whether a SKILL.md actually changes agent behavior (with/without) and which CLI harness does a task best — agent-native subagents or CLI. agentskills.io-compatible. Site: https://cskwork.github.io/skill-ab-eval/
hzang12345-ship-it/hermes-swarm-benchmark Skill
0★Concurrent-agent benchmark suite packaged as a Hermes skill. Markdown REPORT.md output
claudette-agent/agent-eval-kit Skill
no license0★Free mini-eval and scorecard for AI coding agents: Claude Code, Codex, Cursor, Copilot, Windsurf.
yashpatil582/agent-skills/clinical-note-eval Skill
0★ repoAnthropic Agent Skills packaging shipped eval methods: gaming-resistant clinical-note grounding eval + token-F1 scorer. Stdlib-only, runs offline.
XyraSinclair/ideonomy/skills/audit-the-oracle-coverage Skill
0★ repoComputational ideonomy: Gunkel's science of ideas as inference-time machinery — a 37-primitive organon, an MDL-ratcheted respiratory engine, and 14 gated agent skills (Claude Code plugin)
libenxier-beep/codex-custom-skills/skills/autoresearch Skill
0★ repoProduction-grade Codex skills with explicit triggers, deterministic validation, and reusable AI agent workflows.
anupam-io/sprint-skills/skills/sprint-eval Skill
0★ repoAgent skills that turn a one-line goal into a dependency-ordered sprint of GitHub issues, then run them one at a time — review-each (HITL) or fully unattended (AFK). No cluster, no Linear: just gh, claude, and your repo.
satishTheLegend/eval-harness-architect Skill
0★Stand up a rigorous, regression-proof evaluation harness for any LLM/agent system from zero — datasets, scorers, CI gates, and drift monitoring.
satishTheLegend/ai-rmf-governor Skill
0★Claude Code skill: run NIST AI RMF (Govern/Map/Measure/Manage) on your LLM/ML feature — named failure modes, metric thresholds, an offline eval gate, and signed-off residual risk.
hamidettefagh/incident-to-eval Skill
0★A Claude skill that turns a production AI agent incident into a golden eval case. The loop that feeds the two gates.
Paldom/landing-page-builder-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for building high-converting, modern, performant landing pages: design tokens, conversion copywriting, marketing UX, layout patterns, Next.js/React + shadcn implementation, scroll motion, and performance/SEO/accessibility.
Paldom/llm-council-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for building and operating LLM councils - multi-model deliberation with anonymized peer review, robust aggregation, calibrated escalation, and defenses against correlated failure.
StaryMoon/image-benchmark-report-skill/skills/image-benchmark-report Skill
0★ repoAgent Skill for PSNR/SSIM image benchmarks with per-image scores and worst-case reports.
soltraveler-sri/Steerable-Protocol/skills/eval-authoring Skill
0★ repoAn open standard for making apps steerable, shipped with a reference runtime and a skill that lets a coding agent implement it in your codebase.
Nyota-tree/pm-trace-eval/skills/pm-trace-eval Skill
0★ repoPM-facing AI conversation trace eval skill: north-star metrics, 1-5 scores, red/yellow/green, per-turn quote review
Paldom/mcp-server-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for designing and implementing Model Context Protocol (MCP) servers - server architecture, tools, resources, prompts, transports, authorization, testing, and deployment.
Paldom/screenshooter/.claude/skills/add-skill Skill
0★ repoSkills for creating polished screen-recording walkthrough and tutorial videos of web apps, with smooth scripted mouse movements, zoom effects, and explanation cards.
Paldom/stealth-browser-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for stealthy, human-like browser automation with Playwright and Camoufox — real-browser agentic browsing that evades bot detection.
Paldom/promptimize/.claude/skills/add-skill Skill
0★ repoAgent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.
yuanGao0816/prompt-rail Skill
no license0★Measured prompt iteration with train/test dual-gate scoring and anti-overfit rails. Agent Skill (SKILL.md).
avnath13/evalpilot Skill
0★Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.
Paldom/terminaltor/.claude/skills/add-skill Skill
0★ repoAgent Skills for recording, sanitizing, and rendering terminal session demos - capture real command walkthroughs with agents, redact sensitive output, and render polished casts for READMEs and docs.
Paldom/noslop/.claude/skills/add-skill Skill
0★ repoAgent Skills that strip AI writing tells from copy - em dashes, 'it's not X, it's Y' constructions, hedging, and bloat - so text reads like a human wrote it, with evals to verify.
blakebauman/skillist-validator/skills/skill-description-tuning Skill
0★ repoValidate AI coding assistant skills against the Agent Skills specification (agentskills.io) — a dependency-free validator with fix hints, plus skills for auditing instruction quality and tuning description triggering.
megandmartin/agent-skills-repo/skills/agent-mastery/agent-eval-harness Skill
0★ repo75 production-grade agent skills for Hermes Agent + Paperclip — research, write, organize, earn, and run an AI workforce. Every skill passes a QA gate with hard safety rails. Built by Gen AI Hub.
Jorgut/context-quality-suite/skills/eval-harness Skill
0★ repoCaravaca-Labs/puzzletide-cli/skills/puzzletide-agent-evals Skill
0★ repoWord search generator, crossword generator, and sudoku generator & solver in one local-first CLI — printable PDF puzzle worksheets, themed word banks, verifiable LLM evals, and agent skills. From the makers of puzzletide.com.
clauxel/gemini-upgrade-qa-mcp/com.clauxel.geminiupgradeqa/geminiupgradeqa-mcp MCP server
0★ repoRemote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
dbsectrainer/mcp-eval-runner/io.github.dbsectrainer/mcp-eval-runner MCP server
0★ repoA standardized testing harness for MCP servers and agent workflows
lazymac2x/ai-eval-api/io.github.lazymac2x/ai-eval MCP server
no license0★ repoCloudflare Workers MCP server: ai-eval
CreminiAI/skillpack/templates/builtin-skills/skill-creator Skill
1,197★ repoPack and deploy local AI agents for your team in minutes
AvdLee/Xcode-Build-Optimization-Agent-Skill/skills/xcode-build-benchmark Skill
1,195★ repoAn Agent Skill helping you to optimize Xcode incremental and clean builds by running benchmarks and optimizing build settings.
growthack88/growth-marketing-os/skills/benchmark-analyst Skill
79★ repoGrowth Marketing OS | Mahmoud Omar — open-source AI marketing prompts, Claude skills, agents & growth playbooks (EN + AR)
liqiongyu/lenny_skills_plus/skills/ai-evals Skill
51★ repo86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.
tonyblu331/research-proof/skills/research-proof Skill
45★ repoPressure-test research claims with falsifiable evidence plans, adversarial checks, frozen verifiers, and proof ledgers.
NoesisVision/nasde-toolkit/.claude/skills/nasde-benchmark-creator Skill
11★ repoCLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
alexastrum/skl/.agents/skills/golang-benchmark Skill
11★ repoA lightweight, single-binary Agent Skills CLI manager written in Go
di37/EvalSurfer/.cursor/skills/eval-surfer Skill
11★ repoSkill-first, agent-native evaluation protocol for AI apps
Fuenfgeld/pydantic-ai-skills/skills/pydantic-evals Skill
10★ repoProduction-ready Claude Code skills for building AI agents with Pydantic AI. Includes dependency injection, tools, validators, streaming, multi-agent orchestration, and evaluation framework patterns.
yeaight7/agent-powerups/plugins/agent-evaluation-lab/skills/red-team-eval-authoring Skill
6★ repoCurated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Mozurok/fhorja.dev/.claude/skills/ai-feature-eval-harness Skill
6★ repoA workflow operating system for AI-assisted engineering. Task state, decisions, and plans live on disk as files, not in chat history, so context survives across sessions, tools, and restarts.
timurgaleev/vibestack/skills/benchmark-models Skill
5★ repovibestack is a portable skill pack for AI coding agents. Slash commands like /office-hours, /ship, /investigate, /tdd, /review install once and work across every agent that supports the Agent Skills open standard — Claude Code, Cursor, Kiro, and a growing list of others.
josephneumann/claude-corps/skills/benchmark Skill
4★ repoProduction workflow framework for Claude Code — multi-agent orchestration, parallel worktrees, autonomous task execution, and compound engineering skills
build-with-dhiraj/ai-workflow-framework-portability-kit/Plugins/vercel-marketplace-source/.claude/skills/benchmark-e2e Skill
4★ repoPortable, self-contained snapshot of a complete Claude Code setup — 36 specialist agents, 134 skills, plugins, MCP servers & host tooling. Clone, claude login, run one script, restore the whole orchestration stack in ~20 min.
AKSoftCode/aicrew/codex-skills/aicrew-benchmark Skill
3★ repoTDD-first AI dev pipeline for Claude Code, Cursor, Codex, Gemini CLI & Antigravity. /dev /fix /quick + token-saving hooks. npx aicrew install
floomhq/starter/skills/floom-skill-evals Skill
3★ repoFloom Starter Pack. 65 hand-picked AI agent skills for Claude Code, Codex, Cursor, Kimi, OpenCode. One install, auto-activates. +16.2pp pass-rate lift.