agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Eval harness

    uzysjung/uzys-agent-harness/templates/skills/eval-harness Skill

    3 repo

    Curate vetted AI-coding skills & plugins by your tech stack — install only what you need, across Claude Code, Codex, OpenCode & Antigravity

  • Skill creator plus

    yaniv-golan/skill-creator-plus/skill-creator-plus/skills/skill-creator-plus Skill

    3 repo

    Based on Anthropic's skill-creator — with bug fixes, working Cowork support, and official best practices baked in. Claude skill for creating, testing, and improving other Claude skills.

  • Eval rubric generator

    Abhillashjadhav/AI-PM-essential-skills/eval-rubric-generator Skill

    no license2 repo

    Installable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.

  • Eval engine

    Abhillashjadhav/AI-PM-essential-skills/pm-verifier/skills/eval-engine Skill

    no license2 repo

    Installable Claude Code plugins for AI product evaluation, model routing, guarded loops, and MCP migration decisions.

  • Agent optimization

    abhishekgahlot2/agent-optimization/skills/agent-optimization Skill

    2 repo

    Prompt, checklist, and Agent Skill for optimizing LLM agents without benchmark hacks.

  • Bad evals

    ai-creed/ai-shakespii/tests/fixtures/harness/bad-evals Skill

    27 days old2 repo

    Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

  • No evals

    ai-creed/ai-shakespii/tests/fixtures/harness/no-evals Skill

    27 days old2 repo

    Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

  • Four layer eval cascade

    AnthonyAlcaraz/agentic-graph-rag-skills/skills/self-evolution/four-layer-eval-cascade Skill

    2 repo

    Companion repo for Agentic Graph RAG (O'Reilly, Anthony Alcaraz & Sam Julien) — 50 runnable skills + 8 pedagogical notebooks covering all eight chapters, on one moto-mocked AWS DevOps scenario

  • Eval harness

    AtulPurohit/Antigravity-Awesome-Skills/skills/eval-harness Skill

    26 days old2 repo

    Installable GitHub library of 300+ professional agentic skills for Claude Code, Antigravity IDE, Gemini CLI, Cursor, and Copilot. Features a custom NPX installer, 9 stack-specific bundles, validation schemas, security auditing, and an interactive catalog explorer app.

  • Hallucination detector

    AtulPurohit/Antigravity-Awesome-Skills/skills/hallucination-detector Skill

    26 days old2 repo

    Installable GitHub library of 300+ professional agentic skills for Claude Code, Antigravity IDE, Gemini CLI, Cursor, and Copilot. Features a custom NPX installer, 9 stack-specific bundles, validation schemas, security auditing, and an interactive catalog explorer app.

  • Eval harness

    AtulPurohit/Antigravity-Awesome-Skills/plugins/ai-llm-engineering/skills/eval-harness Skill

    26 days old2 repo

    Installable GitHub library of 300+ professional agentic skills for Claude Code, Antigravity IDE, Gemini CLI, Cursor, and Copilot. Features a custom NPX installer, 9 stack-specific bundles, validation schemas, security auditing, and an interactive catalog explorer app.

  • Hallucination detector

    AtulPurohit/Antigravity-Awesome-Skills/plugins/ai-llm-engineering/skills/hallucination-detector Skill

    26 days old2 repo

    Installable GitHub library of 300+ professional agentic skills for Claude Code, Antigravity IDE, Gemini CLI, Cursor, and Copilot. Features a custom NPX installer, 9 stack-specific bundles, validation schemas, security auditing, and an interactive catalog explorer app.

  • Eval skills

    avivsinai/skills-marketplace/plugins/skill-authoring/skills/eval-skills Skill

    2 repo

    Central plugin marketplace for Claude Code and Codex

  • Firm pdca eval

    b2bforce/b2bforce/.agents/skills/firm-pdca-eval Skill

    2 repo

    Skills + workspace for AI agents in B2B service firms

  • Workflow skill evals

    bitwise-media-group/skills/plugins/workflow/skills/workflow-skill-evals Skill

    2 repo

    Coding agent marketplace for skills used by the BitWise Media Group.

  • Agent eval framework

    BuilderCed/agent-skills/skills/eval/agent-eval-framework Skill

    2 repo

    31 cross-platform AI agent skills for regulated industries & underserved markets. EU compliance (AI Act, NIS2, DORA, GDPR), French professional (accounting, tax, notary, real estate), security audit, agent evaluation, Africa mobile money, offline-first.

  • Puzzletide agent evals

    Caravaca-Labs/puzzletide-cli/skills/puzzletide-agent-evals Skill

    20 days old2 repo

    Word search generator, crossword generator, and sudoku generator & solver in one local-first CLI — printable PDF puzzle worksheets, themed word banks, verifiable LLM evals, and agent skills. From the makers of puzzletide.com.

  • Benchmark

    CODE-SAURABH/OpenSkills/benchmark Skill

    11 days old2 repo

    49 production-grade AI agent skills (SKILL.md) for Claude Code, Codex & Antigravity — system design, DevOps, security, QA, and more. MIT licensed, open source.

  • Eval review

    easyinplay/harnessed/workflows/verify/eval-review Skill

    2 repo

    AI coding harness composition orchestrator — manifest-described upstreams, composition skill workflows. Apache-2.0.

  • Skill creator

    furkangonel/cowrangler/bundled_skills/skill-creator Skill

    2 repo

    Autonomous terminal AI agent for workflows and feasible project procedures. Co-Worker Co-Wrangler 🐙

  • Skill creator

    harshitsinghbhandari/domain-expansion/.agents/skills/skill-creator Skill

    no license2 repo

    Collection of Claude agent skills (code audits, LLM councils, PR review, resume tooling) installable individually via npx skills add.

  • Imlazy eval

    hnikoloski/imlazy/skills/imlazy-eval Skill

    26 days old2 repo

    Token-efficient tier-routed agent orchestrator 40+ skills, 19 agents, Obsidian vault. Works with Claude Code, OpenCode, Codex CLI.

  • Lm evaluation harness

    john-data-chen/hermes-agent-backup/skills/mlops/evaluation/lm-evaluation-harness Skill

    2 repo

    backup of hermes-agent

  • Codex skill creator

    jtsang4/efficient-coding/skills/codex-skill-creator Skill

    2 repo

    A curated collection of reusable AI coding skills, MCP server configs, and engineering playbooks for faster, more systematic software development.

  • Eval harness design

    jukrap/ai-agent-playbook/skills/delivery/eval-harness-design Skill

    2 repo

    Reusable AI agent skills, project templates, and guardrails for safer software maintenance and delivery.

  • Junyi xhs benchmark

    junyifei/junyi-skills/junyi-xhs-benchmark Skill

    no license2 repo

    育儿先育己,让 AI 学会你。写给创业者父母的家庭成长 Skills,把你的真实经验和判断标准,交给一个长期服务这个家庭的 Agent。

  • Hugging face community evals

    LiHongwei-cn/lihongwei-cn/mundo-cloud/skills/uncategorized/hugging-face-community-evals Skill

    2 repo

    MUNDO - THE EMPEROR. Complete AI orchestration system with 1208 skills, 25 capability modules, self-evolving, collective consciousness. GitHub Actions 24/7 automation.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/skills/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/anthropic/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Ai incident response desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/ai-incident-response-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval design desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-design-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Eval run analysis desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-run-analysis-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Red team eval desk

    MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/red-team-eval-desk Skill

    2 repo

    Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

  • Agentic eval

    MarieLynneBlock/arcanum-artifex/skills/agentic/agentic-eval Skill

    no license2 repo

    Prompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻‍♀️

  • Eval driven dev

    MarieLynneBlock/arcanum-artifex/skills/agentic/evaluation/eval-driven-dev Skill

    no license2 repo

    Prompts, skills, and agents that survive contact with real workflows. No vendor loyalty. Occasionally heretical. 🧙🏻‍♀️

  • Uncertainty and hallucination

    MedocMay/ai-native-builder-consultant-skills/ai-native-builder-consultant-skills/skills/mental-models/uncertainty-and-hallucination Skill

    no license2 repo

    Consultant-grade skills, workflow, and templates for designing, evaluating, launching, and iterating AI Native products.

  • Agent eval framework

    MedocMay/ai-native-builder-consultant-skills/ai-native-builder-consultant-skills/skills/build-skills/agent-eval-framework Skill

    no license2 repo

    Consultant-grade skills, workflow, and templates for designing, evaluating, launching, and iterating AI Native products.

  • Eval harness

    multiplex-ai/muggle-ai-teams/skills/eval-harness Skill

    2 repo

    AI workflow for Claude Code — describe what you want, get production-grade results. Code, content, design, planning. Built by MuggleTest.

  • Skill creator

    nbialk/quiver-cli/template/.agents/skills/skill-creator Skill

    2 repo

    Compose skills, slash commands & MCP servers from a central catalog into any repo as native configs for opencode, Claude Code and Codex - with lockfile-based drift detection.

  • Radar self eval

    Neetx/ai-research-radar/.claude/skills/radar-self-eval Skill

    no license2 repo

    AI Research Radar — Evidence-Grounded Trend Ledger for AI Systems

  • Eval

    OpenSIN-AI/OpenSIN-Skills/engineering/agenthub/skills/eval Skill

    no license2 repo

    The world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.

  • Self eval

    OpenSIN-AI/OpenSIN-Skills/engineering/self-eval Skill

    no license2 repo

    The world's largest open-source AI agent skill library — 280+ skills across 9 knowledge domains + 46 operational actions. Built on alirezarezvani/claude-skills + OpenSIN-AI.

  • Skill creator

    Poorgramer-Zack/copilot-cli-things/plugins/skill-creator/skills/skill-creator Skill

    2 repo

    A curated collection of extensions, skills, and plugins for GitHub Copilot CLI.

  • Skill creator

    Pyfagorass/bookofspells/skills/anthropic/skill-creator Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Arize evaluator

    Pyfagorass/bookofspells/skills/githubcopilot/arize-evaluator Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Phoenix evals

    Pyfagorass/bookofspells/skills/githubcopilot/phoenix-evals Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Agent platform eval flywheel

    Pyfagorass/bookofspells/skills/google/agent-platform-eval-flywheel Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Huggingface community evals

    Pyfagorass/bookofspells/skills/huggingface/huggingface-community-evals Skill

    no license2 repo

    📖 The Book of Spells: a curated, enchanted index of real LLM tooling — and a pipeline that gathers SKILL.md skills from many houses into one searchable shelf.

  • Skill creator

    rakibulism/agent-skills-os/skills/skill-creator Skill

    2 repo

    THE UNIVERSAL AGENT SKILLS LIBRARY

  • Skill creator

    riggyz/skills/skills/skill-creator Skill

    no license2 repo

    Current skills I am testing out in my agent setup