Evals & benchmarks
944 rows, by stacks then stars
508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid
YousefNabil-SOC/claude-apex/skills/gs-benchmark Skill
2★ repoPublic mirror + teaching install of my Claude Code environment: 1,276 skills, 185 agents, 235 commands, 8 wired MCP servers, three-layer auto-routing (10 CARL domains). Curated original core + the open plugin/marketplace ecosystem, all credited. MIT.
multiplex-ai/muggle-ai-teams/skills/eval-harness Skill
2★ repoAI workflow for Claude Code — describe what you want, get production-grade results. Code, content, design, planning. Built by MuggleTest.
citedy/skills/skills/skill-quality-eval Skill
2★ repoCurated collection of skills by Citedy — AI-powered SEO content automation
patrickserrano/lacquer/profiles/ios/skills/xcode-build-benchmark Skill
2★ repoGo CLI + profile templates that standardize how Claude Code works across every project
OpenSIN-AI/OpenSIN-Skills/engineering/agenthub/skills/eval Skill
no license2★ repoThe world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.
OpenSIN-AI/OpenSIN-Skills/engineering/self-eval Skill
no license2★ repoThe world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.
MarieLynneBlock/arcanum-artifex/skills/agentic/agentic-eval Skill
no license2★ repoPrompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻♀️
MarieLynneBlock/arcanum-artifex/skills/agentic/evaluation/eval-driven-dev Skill
no license2★ repoPrompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻♀️
easyinplay/harnessed/workflows/verify/eval-review Skill
2★ repoAI coding harness composition orchestrator — manifest-described upstreams, composition skill workflows. Apache-2.0.
siddharthrath1999/Claude-Opus-4.7/skills/skill-creator Skill
2★ repoA Claude Opus 4.7 Thinking System that gives advanced reasoning abilities to any other AI, LLM, or Agentic System.
tony/ai-workflow-plugins/.agents/skills/pytest-optimizer-01-benchmark Skill
2★ repoClaude Code Plugins, Commands, and Skills
Ghosteken/agent-harness/archive/skills-community/hugging-face-community-evals Skill
2★ repoProduction-grade engineering skills and specialist agent personas for AI coding assistants. Orchestrates the full SDLC, from spec to ship, with structured verification gates, anti-rationalization checks, and senior-level engineering discipline.
davendra/uk-legal-skills/skills/legal-benchmark Skill
no license2★ repoOpen-source Claude Code skills for England & Wales legal work — 38 /legal commands, 12 agents, and UK legislation + case law MCP servers.
avivsinai/skills-marketplace/plugins/skill-authoring/skills/eval-skills Skill
2★ repoCentral plugin marketplace for Claude Code and Codex
Poorgramer-Zack/copilot-cli-things/plugins/skill-creator/skills/skill-creator Skill
2★ repoA curated collection of extensions, skills, and plugins for GitHub Copilot CLI.
jtsang4/efficient-coding/skills/codex-skill-creator Skill
2★ repoA curated collection of reusable AI coding skills, MCP server configs, and engineering playbooks for faster, more systematic software development.
Abhillashjadhav/AI-PM-essential-skills/eval-rubric-generator Skill
2★ repoInstallable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.
Abhillashjadhav/AI-PM-essential-skills/pm-verifier/skills/eval-engine Skill
2★ repoInstallable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.
rakibulism/agent-skills-os/skills/skill-creator Skill
2★ repoTHE UNIVERSAL AGENT SKILLS LIBRARY
hnikoloski/imlazy/skills/imlazy-eval Skill
2★ repoToken-efficient tier-routed agent orchestrator 40+ skills, 19 agents, Obsidian vault. Works with Claude Code, OpenCode, Codex CLI.
jukrap/ai-agent-playbook/skills/delivery/eval-harness-design Skill
2★ repoReusable AI agent skills, project templates, and guardrails for safer software maintenance and delivery.
sam6dvpte34/social-media-skill/skills/social-engagement-benchmark Skill
no license2★ repoAgent-ready social media crawling and task automation Skills, best for Instagram, TikTok, YouTube, X/Twitter, LinkedIn, Facebook, Reddit, and Xiaohongshu
b2bforce/b2bforce/.agents/skills/firm-pdca-eval Skill
2★ repoSkills + workspace for AI agents in B2B service firms
john-data-chen/hermes-agent-backup/skills/mlops/evaluation/lm-evaluation-harness Skill
2★ repobackup of hermes-agent
furkangonel/cowrangler/bundled_skills/skill-creator Skill
2★ repoAutonomous terminal AI agent for workflows and feasible project procedures. Co-Worker Co-Wrangler 🐙
ai-creed/ai-shakespii/tests/fixtures/harness/bad-evals Skill
2★ repoWorkbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more
ai-creed/ai-shakespii/tests/fixtures/harness/no-evals Skill
2★ repoWorkbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more
riggyz/skills/skills/skill-creator Skill
no license2★ repoCurrent skills I am testing out in my agent setup
BuilderCed/agent-skills/skills/eval/agent-eval-framework Skill
2★ repo31 cross-platform AI agent skills for regulated industries & underserved markets. EU compliance (AI Act, NIS2, DORA, GDPR), French professional (accounting, tax, notary, real estate), security audit, agent evaluation, Africa mobile money, offline-first.
junyifei/junyi-skills/junyi-xhs-benchmark Skill
no license2★ repo育儿先育己,让 AI 学会你。写给创业者父母的家庭成长 Skills,把你的真实经验和判断标准,交给一个长期服务这个家庭的 Agent。
Pyfagorass/bookofspells/skills/anthropic/skill-creator Skill
no license2★ repo📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.
Pyfagorass/bookofspells/skills/githubcopilot/arize-evaluator Skill
no license2★ repo📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.
Pyfagorass/bookofspells/skills/githubcopilot/phoenix-evals Skill
no license2★ repo📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.
Pyfagorass/bookofspells/skills/google/agent-platform-eval-flywheel Skill
no license2★ repo📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.
Pyfagorass/bookofspells/skills/huggingface/huggingface-community-evals Skill
no license2★ repo📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.
MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/ai-incident-response-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-design-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-run-analysis-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/red-team-eval-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/ai-incident-response-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-design-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-run-analysis-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/red-team-eval-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/ai-incident-response-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-design-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-run-analysis-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/red-team-eval-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/ai-incident-response-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-design-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-run-analysis-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/red-team-eval-desk Skill
2★ repoVendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
stephen94125/peon-lib/.skills/skill-creator Skill
2★ repoA lightweight, secure executor for Claude Agent Skills. Inspired by OpenClaw but strictly designed with zero-trust whitelisting, written in Rust.
turntuptechnologies-ai/skills/skills/library-eval Skill
1★ repoTurnt Up Technologies株式会社の Claude Code プラグインマーケットプレイス
Yesterday-AI/ytstack/vendor/gstack/benchmark-models Skill
1★ repoAn opinionated OS for AI coding agents. Plan like a PM, execute like a senior eng. -- Claude Code plugin
Yesterday-AI/ytstack/vendor/gstack/benchmark Skill
1★ repoAn opinionated OS for AI coding agents. Plan like a PM, execute like a senior eng. -- Claude Code plugin
timdevai/proteus/skills/community/from-alireza/c-level-advisor/c-level-agents/skills/cross-eval Skill
1★ repoAlways-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.
timdevai/proteus/skills/community/from-alireza/engineering/agenthub/skills/eval Skill
1★ repoAlways-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.
timdevai/proteus/skills/community/from-alireza/engineering/skills/self-eval Skill
1★ repoAlways-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.
nainishshafi/developer-productivity-skills/.github/skills/llm-as-judge Skill
1★ repoReusable AI skills for Claude Code and GitHub Copilot — scan READMEs, sync forks, create skills, lint Python, and scan for security vulnerabilities
iabhisekbosepm/claude-god-setup/.claude/skills/eval Skill
no license1★ repoMulti-agent orchestration system with 21 specialized agents, automated workflows, and cost-optimized model routing.