agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Hk skill creator

    deepklarity/harness-kit/.claude/skills/hk-skill-creator Skill

    86 repo

    A kit for building with AI agents and also the engineering patterns around it.

  • Agent eval

    mturac/everything-openai-codex/skills/agent-eval Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Benchmark

    mturac/everything-openai-codex/skills/benchmark Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Agent eval

    mturac/everything-openai-codex/docs/ja-JP/skills/agent-eval Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Benchmark

    mturac/everything-openai-codex/docs/ja-JP/skills/benchmark Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/docs/ja-JP/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Healthcare eval harness

    mturac/everything-openai-codex/docs/ja-JP/skills/healthcare-eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/docs/ko-KR/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/docs/tr/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Agent eval

    mturac/everything-openai-codex/docs/zh-CN/skills/agent-eval Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Benchmark

    mturac/everything-openai-codex/docs/zh-CN/skills/benchmark Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/docs/zh-CN/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Healthcare eval harness

    mturac/everything-openai-codex/docs/zh-CN/skills/healthcare-eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/docs/zh-TW/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    mturac/everything-openai-codex/skills/eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Healthcare eval harness

    mturac/everything-openai-codex/skills/healthcare-eval-harness Skill

    82 repo

    EOC: open-source operating system for OpenAI Codex workflows with agents, skills, hooks, rules, memory, safety gates, and cross-harness adapters.

  • Eval harness

    aAAaqwq/AGI-Super-Team/skills/eval-harness Skill

    81 repo

    14 AI executives powered by legendary minds (Musk/Buffett/Simons/Feynman) — deploy your virtual C-Suite in one git clone.

  • Claude code skill creator

    zby/commonplace/kb/work/skill-creator-distillation/sources/claude-code-skill-creator Skill

    81 repo

    The theory of LLM wikis, running as one. A framework for agent-operated knowledge: typed, linked, review-gated markdown your agents execute.

  • Llm benchmark

    KerberosClaw/kc_ai_skills/llm-benchmark Skill

    75 repo

    AI Skills That Actually Do Things — 中文優先的 Claude Code / Codex agent skills 合集 · Reusable bilingual skills for any LLM workflow

  • Benchmark analyst

    growthack88/growth-marketing-os/skills/benchmark-analyst Skill

    66 repo

    Growth Marketing OS | Mahmoud Omar — open-source AI marketing prompts, Claude skills, agents & growth playbooks (EN + AR)

  • Hallucination prevention

    DevelopersGlobal/ai-agent-skills/skills/hallucination-prevention Skill

    64 repo

    AI agent skills for production grade applications

  • Decision eval

    avelikiy/great_cto/skills/decision-eval Skill

    60 repo

    Don't buy software. Get the work done. GreatCTO ships AI autopilots that run a whole business function — medical coding, legal docs, procurement, accounting, IT, tax — from intake to outcome. A qualified human signs only the judgment calls. Live connectors, built-in compliance.

  • Genesis evals

    danielmeppiel/genesis/dev/skills/genesis-evals Skill

    60 repo

    Markdown that steers an LLM is code. Genesis is the architectural layer for designing multi-agent, multi-skill systems -- with named patterns, contracts, and substrate portability, before you write them.

  • Agent eval

    ComeOnOliver/skillshub/skills/affaan-m/everything-claude-code/agent-eval Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Benchmark

    ComeOnOliver/skillshub/skills/affaan-m/everything-claude-code/benchmark Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Eval harness

    ComeOnOliver/skillshub/skills/affaan-m/everything-claude-code/eval-harness Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Benchmark kernel

    ComeOnOliver/skillshub/skills/aiskillstore/marketplace/flashinfer-ai/benchmark-kernel Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Eval

    ComeOnOliver/skillshub/skills/alirezarezvani/claude-skills/eval Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill creator

    ComeOnOliver/skillshub/skills/anthropics/skills/skill-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Ue benchmark

    ComeOnOliver/skillshub/skills/blackplume233/UnrealMCPHub/ue-benchmark Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Benchmark email automation

    ComeOnOliver/skillshub/skills/ComposioHQ/awesome-claude-skills/benchmark-email-automation Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Competitive feature benchmark

    ComeOnOliver/skillshub/skills/comsky/remy-skill-recipes/competitive-feature-benchmark Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill creator

    ComeOnOliver/skillshub/skills/countbot-ai/CountBot/skill-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Promptfoo evaluation

    ComeOnOliver/skillshub/skills/daymade/claude-code-skills/promptfoo-evaluation Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill creator

    ComeOnOliver/skillshub/skills/daymade/claude-code-skills/skill-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Agentic eval

    ComeOnOliver/skillshub/skills/github/awesome-copilot/agentic-eval Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Eval driven dev

    ComeOnOliver/skillshub/skills/github/awesome-copilot/eval-driven-dev Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill creator

    ComeOnOliver/skillshub/skills/guanyang/antigravity-skills/skill-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Golang benchmark

    ComeOnOliver/skillshub/skills/Harmeet10000/skills/golang-benchmark Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill creator

    ComeOnOliver/skillshub/skills/Harmeet10000/skills/skill-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Benchmark suite creator

    ComeOnOliver/skillshub/skills/jeremylongshore/claude-code-plugins-plus-skills/benchmark-suite-creator Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Security benchmark runner

    ComeOnOliver/skillshub/skills/jeremylongshore/claude-code-plugins-plus-skills/security-benchmark-runner Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Agentic eval

    ComeOnOliver/skillshub/skills/LeoYeAI/openclaw-master-skills/agentic-eval Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Bigcode evaluation harness

    ComeOnOliver/skillshub/skills/Orchestra-Research/AI-Research-SKILLs/bigcode-evaluation-harness Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Lm evaluation harness

    ComeOnOliver/skillshub/skills/Orchestra-Research/AI-Research-SKILLs/lm-evaluation-harness Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Ai eval ci

    ComeOnOliver/skillshub/skills/TerminalSkills/skills/ai-eval-ci Skill

    59 repo

    🧠 The right skill, one API call. AI agent skills registry with token-efficient skill resolution. 5,000+ skills from 500+ top repos.

  • Skill evaluator

    HeshamFS/materials-simulation-skills/skills/meta/skill-evaluator Skill

    58 repo

    Agent Skills for computational materials science -- numerical stability, solvers, meshing, convergence, and simulation workflows.

  • Benchmark and mms planner

    HeshamFS/materials-simulation-skills/skills/verification-validation/benchmark-and-mms-planner Skill

    58 repo

    Agent Skills for computational materials science -- numerical stability, solvers, meshing, convergence, and simulation workflows.

  • Alterlab scholar eval

    AlterLab-IEU/AlterLab-Academic-Skills/skills/writing-tools/alterlab-scholar-eval Skill

    51 repo

    239 evaluated academic Claude/agent skills across 17 research domains (bioinformatics, data science, clinical, social-science methods, Turkish academia & more). Executable eval per skill, deterministic citation verifier, research→write→review→publish pipeline, and a skill-finder front door. Claude Code, Cursor, Codex, Gemini CLI & Copilot.

  • Ai evals

    liqiongyu/lenny_skills_plus/skills/ai-evals Skill

    51 repo

    86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.

  • Building with llms

    liqiongyu/lenny_skills_plus/skills/building-with-llms Skill

    51 repo

    86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.

  • Writing prds executable

    liqiongyu/lenny_skills_plus/samples/writing-prds-executable Skill

    51 repo

    86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.

  • Eval

    Kanevry/session-orchestrator/skills/eval Skill

    47 repo

    Loop engineering for AI coding agents — turn ad-hoc sessions into a repeatable research → plan → wave-execute → close loop with verification gates. Runs on Claude Code, Codex CLI, Cursor, and Pi. MIT community plugin.

  • Autoresearch

    adriannoes/awesome-agentic-ai/cursor-claude-codex/skills/autoresearch Skill

    46 repo

    329 agent skills (Cursor, Claude Code & Codex), 5,380 OpenClaw skills, 201 ML notebooks, 7 textbooks, 52 research papers, 17 industry reports for PMs, Designers & Developers.

  • Run deep swe

    adriannoes/awesome-agentic-ai/cursor-claude-codex/skills/david-ondrej/agent-orchestration/run-deep-swe Skill

    46 repo

    329 agent skills (Cursor, Claude Code & Codex), 5,380 OpenClaw skills, 201 ML notebooks, 7 textbooks, 52 research papers, 17 industry reports for PMs, Designers & Developers.

  • Evals

    adriannoes/awesome-agentic-ai/cursor-claude-codex/skills/igoruehara-spec-driven/skills/evals Skill

    46 repo

    329 agent skills (Cursor, Claude Code & Codex), 5,380 OpenClaw skills, 201 ML notebooks, 7 textbooks, 52 research papers, 17 industry reports for PMs, Designers & Developers.

  • Ai system testing

    petrkindlmann/qa-skills/skills/ai-system-testing Skill

    45 repo

    50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime.

  • Creator commercial value benchmark

    allinherog-star/ai-skills/skills/creator-commercial-value-benchmark Skill

    no license43 repo

    AI Skills(ai-skills.ai)是一个面向各行各业的 AI 技能库。可直接用于openclaw,Hermes,qclaw,claude code,codex等智能体平台。为你的职业发展、行业竞争力,增加更多可能。

  • Agent evals

    BagelHole/DevOps-Security-Agent-Skills/devops/ai/agent-evals Skill

    43 repo

    Agent-ready DevOps, security, infrastructure, and compliance knowledge base with 80+ skills across Kubernetes, Terraform, AWS/Azure/GCP, AI platform operations, container hardening, SOC2/ISO27001, and incident response—plus ready-to-run scripts, templates, and playbooks for SRE, platform, and security teams.

  • Rag observability evals

    BagelHole/DevOps-Security-Agent-Skills/devops/ai/rag-observability-evals Skill

    43 repo

    Agent-ready DevOps, security, infrastructure, and compliance knowledge base with 80+ skills across Kubernetes, Terraform, AWS/Azure/GCP, AI platform operations, container hardening, SOC2/ISO27001, and incident response—plus ready-to-run scripts, templates, and playbooks for SRE, platform, and security teams.