agentsclimarketplace

Evals & benchmarks

944 rows, by stacks then stars

508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid

  • Gs benchmark

    YousefNabil-SOC/claude-apex/skills/gs-benchmark Skill

    2 repo

    Public mirror + teaching install of my Claude Code environment: 1,276 skills, 185 agents, 235 commands, 8 wired MCP servers, three-layer auto-routing (10 CARL domains). Curated original core + the open plugin/marketplace ecosystem, all credited. MIT.

  • Eval harness

    multiplex-ai/muggle-ai-teams/skills/eval-harness Skill

    2 repo

    AI workflow for Claude Code — describe what you want, get production-grade results. Code, content, design, planning. Built by MuggleTest.

  • Skill quality eval

    citedy/skills/skills/skill-quality-eval Skill

    2 repo

    Curated collection of skills by Citedy — AI-powered SEO content automation

  • Xcode build benchmark

    patrickserrano/lacquer/profiles/ios/skills/xcode-build-benchmark Skill

    2 repo

    Go CLI + profile templates that standardize how Claude Code works across every project

  • Eval

    OpenSIN-AI/OpenSIN-Skills/engineering/agenthub/skills/eval Skill

    no license2 repo

    The world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.

  • Self eval

    OpenSIN-AI/OpenSIN-Skills/engineering/self-eval Skill

    no license2 repo

    The world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.

  • Agentic eval

    MarieLynneBlock/arcanum-artifex/skills/agentic/agentic-eval Skill

    no license2 repo

    Prompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻‍♀️

  • Eval driven dev

    MarieLynneBlock/arcanum-artifex/skills/agentic/evaluation/eval-driven-dev Skill

    no license2 repo

    Prompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻‍♀️

  • Eval review

    easyinplay/harnessed/workflows/verify/eval-review Skill

    2 repo

    AI coding harness composition orchestrator — manifest-described upstreams, composition skill workflows. Apache-2.0.

  • Skill creator

    siddharthrath1999/Claude-Opus-4.7/skills/skill-creator Skill

    2 repo

    A Claude Opus 4.7 Thinking System that gives advanced reasoning abilities to any other AI, LLM, or Agentic System.

  • Pytest optimizer 01 benchmark

    tony/ai-workflow-plugins/.agents/skills/pytest-optimizer-01-benchmark Skill

    2 repo

    Claude Code Plugins, Commands, and Skills

  • Hugging face community evals

    Ghosteken/agent-harness/archive/skills-community/hugging-face-community-evals Skill

    2 repo

    Production-grade engineering skills and specialist agent personas for AI coding assistants. Orchestrates the full SDLC, from spec to ship, with structured verification gates, anti-rationalization checks, and senior-level engineering discipline.

  • Legal benchmark

    davendra/uk-legal-skills/skills/legal-benchmark Skill

    no license2 repo

    Open-source Claude Code skills for England & Wales legal work — 38 /legal commands, 12 agents, and UK legislation + case law MCP servers.

  • Eval skills

    avivsinai/skills-marketplace/plugins/skill-authoring/skills/eval-skills Skill

    2 repo

    Central plugin marketplace for Claude Code and Codex

  • Skill creator

    Poorgramer-Zack/copilot-cli-things/plugins/skill-creator/skills/skill-creator Skill

    2 repo

    A curated collection of extensions, skills, and plugins for GitHub Copilot CLI.

  • Codex skill creator

    jtsang4/efficient-coding/skills/codex-skill-creator Skill

    2 repo

    A curated collection of reusable AI coding skills, MCP server configs, and engineering playbooks for faster, more systematic software development.

  • Eval rubric generator

    Abhillashjadhav/AI-PM-essential-skills/eval-rubric-generator Skill

    2 repo

    Installable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.

  • Eval engine

    Abhillashjadhav/AI-PM-essential-skills/pm-verifier/skills/eval-engine Skill

    2 repo

    Installable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.

  • Skill creator

    rakibulism/agent-skills-os/skills/skill-creator Skill

    2 repo

    THE UNIVERSAL AGENT SKILLS LIBRARY

  • Imlazy eval

    hnikoloski/imlazy/skills/imlazy-eval Skill

    2 repo

    Token-efficient tier-routed agent orchestrator 40+ skills, 19 agents, Obsidian vault. Works with Claude Code, OpenCode, Codex CLI.

  • Eval harness design

    jukrap/ai-agent-playbook/skills/delivery/eval-harness-design Skill

    2 repo

    Reusable AI agent skills, project templates, and guardrails for safer software maintenance and delivery.

  • Social engagement benchmark

    sam6dvpte34/social-media-skill/skills/social-engagement-benchmark Skill

    no license2 repo

    Agent-ready social media crawling and task automation Skills, best for Instagram, TikTok, YouTube, X/Twitter, LinkedIn, Facebook, Reddit, and Xiaohongshu

  • Firm pdca eval

    b2bforce/b2bforce/.agents/skills/firm-pdca-eval Skill

    2 repo

    Skills + workspace for AI agents in B2B service firms

  • Lm evaluation harness

    john-data-chen/hermes-agent-backup/skills/mlops/evaluation/lm-evaluation-harness Skill

    2 repo

    backup of hermes-agent

  • Skill creator

    furkangonel/cowrangler/bundled_skills/skill-creator Skill

    2 repo

    Autonomous terminal AI agent for workflows and feasible project procedures. Co-Worker Co-Wrangler 🐙

  • Bad evals

    ai-creed/ai-shakespii/tests/fixtures/harness/bad-evals Skill

    2 repo

    Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

  • No evals

    ai-creed/ai-shakespii/tests/fixtures/harness/no-evals Skill

    2 repo

    Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

  • Skill creator

    riggyz/skills/skills/skill-creator Skill

    no license2 repo

    Current skills I am testing out in my agent setup

  • Agent eval framework

    BuilderCed/agent-skills/skills/eval/agent-eval-framework Skill

    2 repo

    31 cross-platform AI agent skills for regulated industries & underserved markets. EU compliance (AI Act, NIS2, DORA, GDPR), French professional (accounting, tax, notary, real estate), security audit, agent evaluation, Africa mobile money, offline-first.

  • Junyi xhs benchmark

    junyifei/junyi-skills/junyi-xhs-benchmark Skill

    no license2 repo

    育儿先育己,让 AI 学会你。写给创业者父母的家庭成长 Skills,把你的真实经验和判断标准,交给一个长期服务这个家庭的 Agent。

  • Skill creator

    Pyfagorass/bookofspells/skills/anthropic/skill-creator Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Arize evaluator

    Pyfagorass/bookofspells/skills/githubcopilot/arize-evaluator Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Phoenix evals

    Pyfagorass/bookofspells/skills/githubcopilot/phoenix-evals Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Agent platform eval flywheel

    Pyfagorass/bookofspells/skills/google/agent-platform-eval-flywheel Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Huggingface community evals

    Pyfagorass/bookofspells/skills/huggingface/huggingface-community-evals Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Skill creator

    stephen94125/peon-lib/.skills/skill-creator Skill

    2 repo

    A lightweight, secure executor for Claude Agent Skills. Inspired by OpenClaw but strictly designed with zero-trust whitelisting, written in Rust.

  • Library eval

    turntuptechnologies-ai/skills/skills/library-eval Skill

    1 repo

    Turnt Up Technologies株式会社の Claude Code プラグインマーケットプレイス

  • Benchmark models

    Yesterday-AI/ytstack/vendor/gstack/benchmark-models Skill

    1 repo

    An opinionated OS for AI coding agents. Plan like a PM, execute like a senior eng. -- Claude Code plugin

  • Benchmark

    Yesterday-AI/ytstack/vendor/gstack/benchmark Skill

    1 repo

    An opinionated OS for AI coding agents. Plan like a PM, execute like a senior eng. -- Claude Code plugin

  • Cross eval

    timdevai/proteus/skills/community/from-alireza/c-level-advisor/c-level-agents/skills/cross-eval Skill

    1 repo

    Always-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.

  • Eval

    timdevai/proteus/skills/community/from-alireza/engineering/agenthub/skills/eval Skill

    1 repo

    Always-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.

  • Self eval

    timdevai/proteus/skills/community/from-alireza/engineering/skills/self-eval Skill

    1 repo

    Always-on multi-agent AI workstation for Claude Code. 142 specialist agents (ML, security, K8s, backend, frontend) · 366 curated skills · 200+ prompt templates · auto prompt-matcher · auto skill-matcher · MCP support · token router cuts Anthropic bill 6x via Kimi/Haiku/Ollama. LiteLLM · BYOK · MIT + Apache. One-line install.

  • Llm as judge

    nainishshafi/developer-productivity-skills/.github/skills/llm-as-judge Skill

    1 repo

    Reusable AI skills for Claude Code and GitHub Copilot — scan READMEs, sync forks, create skills, lint Python, and scan for security vulnerabilities

  • Eval

    iabhisekbosepm/claude-god-setup/.claude/skills/eval Skill

    no license1 repo

    Multi-agent orchestration system with 21 specialized agents, automated workflows, and cost-optimized model routing.