Evals & benchmarks
946 rows, by stacks then stars
hamidettefagh/incident-to-eval Skill
0★A Claude skill that turns a production AI agent incident into a golden eval case. The loop that feeds the two gates.
hamidettefagh/two-gates/skills/incident-to-eval Skill
0★ repoA Claude Code plugin that installs the two-gate method for production AI agents: a design gate, a ship gate, and an incident-to-eval loop that turns failures into regression tests.
Llm eval harness and scoring pipeline
HamzaYM/reliable-ai-skills/skills/evals-and-scoring/llm-eval-harness-and-scoring-pipeline Skill
29 days old0★ repoA curated library of production-tested skills for building reliable AI systems with Claude Code. By Hamza Malik, hamz.ai.
Hefrock/agent-skills/skills/agent-eval Skill
0★ repoPortable AI agent skills built on the open SKILL.md standard — self-contained capabilities any compatible agent can discover and load on demand, usable across Claude, Codex, Gemini CLI, Cursor, and GitHub Copilot.
Hefrock/agent-skills/skills/deid-reid-harness Skill
0★ repoPortable AI agent skills built on the open SKILL.md standard — self-contained capabilities any compatible agent can discover and load on demand, usable across Claude, Codex, Gemini CLI, Cursor, and GitHub Copilot.
hzang12345-ship-it/hermes-swarm-benchmark Skill
0★Concurrent-agent benchmark suite packaged as a Hermes skill. Markdown REPORT.md output
iamdemetris/lude-kit/skills/personas/senior-eval-engineer Skill
no license0★ repoA curated library of senior grade Agent Skills and subagents for Claude Code and OpenAI Codex. 70 skills, 30 dispatchable subagents, designed for multi agent orchestration.
ifBars/blender-agent-studio/plugins/blender-agent-studio/skills/blender-agent-benchmark Skill
12 days old0★ repoCodex plugin for reproducible Blender modeling, validation, animation, MCP tooling, and agent benchmarking
jacob-balslev/skills/skills/ai-engineering/eval-driven-development Skill
0★ repoPublic Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills
jacob-balslev/skills/skills/ai-engineering/prompt-injection-defense Skill
0★ repoPublic Agent Skills library exported from skill-graph. Install: npx skills add jacob-balslev/skills
jae-labs/skills/skills/skill-creator Skill
0★ repoAgent Skills.
Jaypatel1511/cdfi-superpowers/skills/cdfi-peer-benchmark Skill
28 days old0★ repoAI skills for NMTC eligibility, bank-CDFI peer benchmarking & HMDA analysis — grounded in audited PyPI tools, not hallucinated.
Jorgut/context-quality-suite/skills/eval-harness Skill
14 days old0★ repojoshuaporth/appsec-skill/.cursor/skills/benchmark Skill
0★ repoPortable secure code review skill for AI coding agents — OWASP/CWE coverage, structured findings, and remediation guidance. Works with Cursor, Claude Code, Kiro, and Open Agent Skills.
jpoindexter/product-management-skills/skills/product-ai-evals Skill
14 days old0★ repoOperational product-management skills and a /pm router for Codex and Claude.
jpoindexter/reward-seeking-safety-skills/skills/contrastive-authority-eval Skill
15 days old0★ repoAgent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.
jpoindexter/reward-seeking-safety-skills/skills/eval-awareness-red-team Skill
15 days old0★ repoAgent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.
justinramos101/agent-skill-kit/skills/eval-best-practices Skill
0★ repoBattle-tested Agent Skills for coding agents — source-grounded, failure-driven heuristics that audit and design real surfaces. Install with npx skills.
K-9Nine/skipthespec/eval-loops Skill
0★ repoClaude Code skills for prototype-first, autonomous AI development — skip the PRD, dogfood the prototype, run automated eval loops, ship behind flags. Anthropic's playbook as installable Agent Skills.
Kbediako/evergreen-codex-skills/skills/native-agent-evals Skill
21 days old0★ repoEvergreen, generally useful Codex skills designed to stay valuable as models improve.
kennyolofsson23-netizen/claude-code-config/skills/skill-creator Skill
no license0★ repoThe most comprehensive Claude Code setup — 29 skills, 9 plugins, 13 hooks, 8 agents, 8 commands. Self-documenting, self-improving, CLI-first. Includes self-interview prompt for personalization.
kjuhwa/skills-hub/skills/cli/gstack/benchmark Skill
0★ repoSelf-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.
kjuhwa/skills-hub/skills/llm-agents/llm-as-judge-loop Skill
0★ repoSelf-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.
Tts trailing silence and hallucination trim
kjuhwa/skills-hub/skills/ml-ops/tts-trailing-silence-and-hallucination-trim Skill
0★ repoSelf-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.
Threadpool benchmark with warmup
kjuhwa/skills-hub/skills/observability/threadpool-benchmark-with-warmup Skill
0★ repoSelf-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.
kjuhwa/skills-hub/skills/testing/benchmark Skill
0★ repoSelf-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.
knownasnaffy/prompthound/dataset/case_00228 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_00490 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_01169 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_01745 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_02998 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_03442 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_03796 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
knownasnaffy/prompthound/dataset/case_05375 Skill
no license0★ repoA fast, offline static risk analysis CLI for AI agent skill files. Detects malicious instructions, steganographic payloads, and dangerous capability chains.
kochellenk-afk/google-ads-diagnostic-toolkit/skills/competitor-benchmark-report Skill
19 days old0★ repo10 production Claude Skills covering the full Google Ads diagnostic lifecycle: waste, spikes, Quality Score, budgets, reporting
lawwu/skills-marketplace/{{cookiecutter.repo_name}}/plugins/{{cookiecutter.plugin_name}}/skills/skill-creator Skill
no license0★ repoCookiecutter Template for an Agent Skills Marketplace
lazymac2x/ai-eval-api/io.github.lazymac2x/ai-eval MCP server
no license0★ repoCloudflare Workers MCP server: ai-eval
learn-with-santosh/claude-master-skills/skills/skill-creator Skill
no license0★ repoA curated collection of specialized skills and workflows designed to enhance Claude's capabilities in specific domains. These skills provide frameworks, psychological triggers, and structured processes to deliver high-quality, professional results.
Levix0501/next-agent-rails/.agents/skills/skill-creator Skill
0★ repoOpinionated Next.js 16 starter for building with AI agents — ship features fast while architecture and directory structure stay under control: fixed structure, enforced conventions, curated libraries, preinstalled skills, native toolchain.
lhbsaa/apex-discovery/skills/arbor Skill
15 days old0★ repoScientific research extension for apex-unified - 149 scientific skills for biology, chemistry, medicine, materials research. Powered by Scientific Agent Skills (31.4k star).
libenxier-beep/codex-custom-skills/skills/autoresearch Skill
0★ repoProduction-grade Codex skills with explicit triggers, deterministic validation, and reusable AI agent workflows.
Liber1917/agent-skill-infra/skills/test-runner Skill
no license0★ repoAgent Skill Infrastructure: quality check, behavior test runner, version awareness
LoogacyStudio/skills/.github/skills/benchmark-author-item Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
Benchmark author rerun manifest
LoogacyStudio/skills/.github/skills/benchmark-author-rerun-manifest Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-author-skill-evals Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-author-variants Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-cluster-failures Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-core Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-judge-run Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
LoogacyStudio/skills/.github/skills/benchmark-review-candidate Skill
0★ repoThis repository stores reusable agent skills, repo level benchmark workflow material, and plugin bundles for coding agents.
MarchRory/liushi_agent_settings/.agents/skills/eval-runner Skill
no license0★ repoPortable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.
MarchRory/liushi_agent_settings/packages/harness/src/.agents/skills/eval-runner Skill
no license0★ repoPortable Codex agent harness and config monorepo with reusable roles, skills, validation, eval tooling, and safe project install sync.
marktantongco/opencodelinux/skills/skill-creator Skill
no license0★ repoMerged opencode config + agents + accomplishments showcase — 217 skills, 17 agent profiles, 6 MCP servers, 78-server registry
matematicsolutions/awesome-matematic-skills-en/verification-foundation/skills/output-scoring-en Skill
no license0★ repoEnglish hub of method-neutral legal AI skills (verification core, content quality, EU law) - bundle model. Polish-jurisdiction skills live in awesome-matematic-skills-pl.
MauricioQuezadaHaintech/karvey/plugins/karvey/skills/karvey-benchmark-models Skill
0★ repoKarvey — método spec-driven development agnóstico de stack (Afán, selknam). Plugin de Claude Code. © HainTech, Apache 2.0.
MauroProto/mis-skills/external/anthropics-skills/skill-creator Skill
0★ repoSkills personales de diseño frontend (taste, brutalist, minimalist, soft, etc.) + cross-agent sync para Claude Code, Codex, Cursor, Continue, Antigravity. SKILL.md compatible.
mcmespinaa/thesis-theory-evaluator Skill
0★Claude Code skill that evaluates a thesis Theoretical Framework chapter against the OL646E (Malmö University SALSU) course rubric. Companion to thesis-anchoring-map.
megandmartin/agent-skills-repo/skills/agent-mastery/agent-eval-harness Skill
12 days old0★ repo75 production-grade agent skills for Hermes Agent + Paperclip — research, write, organize, earn, and run an AI workforce. Every skill passes a QA gate with hard safety rails. Built by Gen AI Hub.
Methasit-Pun/ts-ddd-clean-architecture/.claude/skills/agent-eval Skill
0★ repoTypeScript DDD Clean Architecture Skills
michelve/hugin-v0/skills/skill-creator Skill
0★ repoA Claude Code plugin packaging 23 skills, 8 agents, 5 event hooks, and 7 MCP servers for full-stack development with React 19, TypeScript, Express, Prisma, Tailwind CSS v4, and shadcn/ui.