agentsclimarketplace

Evals & benchmarks

944 rows, by stacks then stars

508 of these skill files have been read, by 499 distinct authors, and what they tell an agent counted. What Evals & benchmarks authors agree on, and what they forbid

  • Doca flow perf

    NVIDIA/skills/skills/doca-flow-perf Skill

    2,853 repo

    Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

  • Jetson llm benchmark

    NVIDIA/skills/skills/jetson-llm-benchmark Skill

    2,853 repo

    Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

  • Rag eval

    NVIDIA/skills/skills/rag-eval Skill

    2,853 repo

    Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

  • Evals context

    foryourhealth111-pixel/Vibe-Skills/bundled/skills/evals-context Skill

    2,656 repo

    VibeSkills is a general-purpose Skill that automatically routes local Skills and intelligently orchestrates harness workflows.

  • Cortex eval

    jeremylongshore/claude-code-plugins-plus-skills/plugins/ai-agency/tonone/skills/cortex-eval Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Langchain eval harness

    jeremylongshore/claude-code-plugins-plus-skills/plugins/saas-packs/langchain-py-pack/skills/langchain-eval-harness Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Detecting eval exec usage

    jeremylongshore/claude-code-plugins-plus-skills/plugins/security/penetration-tester/skills/detecting-eval-exec-usage Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Detecting eval exec usage

    jeremylongshore/claude-code-plugins-plus-skills/skills/.curated/detecting-eval-exec-usage Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Langchain eval harness

    jeremylongshore/claude-code-plugins-plus-skills/skills/.curated/langchain-eval-harness Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Security benchmark runner

    jeremylongshore/claude-code-plugins-plus-skills/skills/04-security-advanced/security-benchmark-runner Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Benchmark suite creator

    jeremylongshore/claude-code-plugins-plus-skills/skills/10-performance-testing/benchmark-suite-creator Skill

    2,628 repo

    425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

  • Eval harness

    a5c-ai/babysitter/library/methodologies/everything-claude-code/skills/eval-harness Skill

    1,665 repo

    Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration

  • Giab benchmark validator

    a5c-ai/babysitter/library/specializations/domains/science/bioinformatics/skills/giab-benchmark-validator Skill

    1,665 repo

    Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration

  • Benchmark suite manager

    a5c-ai/babysitter/library/specializations/domains/science/computer-science/skills/benchmark-suite-manager Skill

    1,665 repo

    Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration

  • Benchmark suite manager

    a5c-ai/babysitter/library/specializations/domains/science/mathematics/skills/benchmark-suite-manager Skill

    1,665 repo

    Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration

  • Performance benchmark suite

    a5c-ai/babysitter/library/specializations/sdk-platform-development/skills/performance-benchmark-suite Skill

    1,665 repo

    Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration

  • Skill creator

    feiskyer/claude-code-settings/skills/skill-creator Skill

    1,625 repo

    Curated skills, sub-agents, and config templates that supercharge Claude Code — research, image gen, GitHub automation & more.

  • Respond to eval

    pedrohcgs/claude-code-my-workflow/.claude/skills/respond-to-eval Skill

    1,465 repo

    A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols.

  • Benchmark

    hoangsonww/Claude-Code-Agent-Monitor/plugins/ccam-insights/skills/benchmark Skill

    891 repo

    🚀 A real-time monitoring dashboard for Claude Code, built with SQLite3, Node.js, Express, React, Vite, TailwindCSS, and WebSockets. It tracks sessions, agent activity, tool usage, and subagent orchestration, providing live analytics, a Kanban status board, status notifications, a cute buddy, and an interactive web UI/MacOS/Windows native app.

  • Model evaluation

    awslabs/agent-plugins/plugins/sagemaker-ai/skills/model-evaluation Skill

    856 repo

    Agent Plugins for AWS equip AI coding agents with the skills to help you architect, deploy, and operate on AWS.

  • 00 meta eval

    agentscope-ai/OpenJudge/skills/eval_pipeline/00-meta-eval Skill

    791 repo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • 01 eval design

    agentscope-ai/OpenJudge/skills/eval_pipeline/01-eval-design Skill

    791 repo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • 04 eval report

    agentscope-ai/OpenJudge/skills/eval_pipeline/04-eval-report Skill

    791 repo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • 05 rag eval

    agentscope-ai/OpenJudge/skills/eval_pipeline/05-rag-eval Skill

    791 repo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • Ref hallucination arena

    agentscope-ai/OpenJudge/skills/ref-hallucination-arena Skill

    791 repo

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

  • Skill creator

    countbot-ai/CountBot/workspace/skills/skill-creator Skill

    762 repo

    更适配中文用户的轻量开源AI Agent | 国产大模型Coding plan支持 | 兼容OpenClaw Skills生态| 已接入微信ClawBot/微博龙虾/飞书/钉钉/QQ/小智AI/Telegram/deepseek-v4。

  • Eval config

    indranilbanerjee/digital-marketing-pro/skills/eval-config Skill

    734 repo

    Open-source AI marketing plugin for agencies & in-house teams — 158 skills, 25 specialist agents, 12-Part Strategy Flow, Cowork team-persistent, EU AI Act Article 50 ready, 6-platform AEO/GEO incl. Google AI Mode. Installs on Claude Code, Cowork, Codex, Cursor, Copilot CLI, Antigravity. MIT-licensed.

  • Eval content

    indranilbanerjee/digital-marketing-pro/skills/eval-content Skill

    734 repo

    Open-source AI marketing plugin for agencies & in-house teams — 158 skills, 25 specialist agents, 12-Part Strategy Flow, Cowork team-persistent, EU AI Act Article 50 ready, 6-platform AEO/GEO incl. Google AI Mode. Installs on Claude Code, Cowork, Codex, Cursor, Copilot CLI, Antigravity. MIT-licensed.

  • Eval suite

    indranilbanerjee/digital-marketing-pro/skills/eval-suite Skill

    734 repo

    Open-source AI marketing plugin for agencies & in-house teams — 158 skills, 25 specialist agents, 12-Part Strategy Flow, Cowork team-persistent, EU AI Act Article 50 ready, 6-platform AEO/GEO incl. Google AI Mode. Installs on Claude Code, Cowork, Codex, Cursor, Copilot CLI, Antigravity. MIT-licensed.

  • Skill creator

    ECNU-ICALK/AutoSkill/SkillBank/Common/anthropics-skill/skill-creator Skill

    no license551 repo

    AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution

  • Ln 34 benchmark comparator

    levnikolaevich/claude-code-skills/plugins/optimization-suite/skills/ln-34-benchmark-comparator Skill

    535 repo

    18 standalone skills for Claude Code and Codex: review, audit, optimization, testing, product discovery, and safe repository publishing.

  • 03 performance eval global

    minhnv0807/ai-business-skills/skills/en/03-performance-eval-global Skill

    527 repo

    63 bilingual AI marketing skills (31 VN + 31 Global) for Claude Code, OpenCode, Codex, VS Code. Marketing strategy, content production, performance analytics, personal brand, AI avatar, dropshipping mastery, design master (8 design types). 4 regions (US/EU/SEA/LATAM) + Vietnam 2025-2026. Anthropic-pattern aligned. Companion: opa-kit.

  • Agent benchmark

    vibeeval/vibecosystem/skills/agent-benchmark Skill

    526 repo

    AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution.

  • Eval harness

    vibeeval/vibecosystem/skills/eval-harness Skill

    526 repo

    AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution.

  • Skill eval

    notque/vexjoy-agent/skills/meta/skill-eval Skill

    415 repo

    VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.

  • Skill creator

    aiskillstore/marketplace/skills/davila7/skill-creator Skill

    no license408 repo

    Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.

  • Skill development

    aiskillstore/marketplace/skills/davila7/skill-development Skill

    no license408 repo

    Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.

  • Benchmark kernel

    aiskillstore/marketplace/skills/flashinfer-ai/benchmark-kernel Skill

    no license408 repo

    Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.

  • Google agents cli eval

    aiskillstore/marketplace/skills/google/google-agents-cli-eval Skill

    no license408 repo

    Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.

  • Eval user

    bruc3van/agent-skills-guard/src-tauri/tests/fixtures/security/rule_matrix/p0-signatures/eval-user Skill

    no license381 repo

    一款提供Agent Skills安全扫描和可视化管理的桌面应用 | A desktop application that provides security scanning and visual management for Agent Skills.

  • Agents skill creator

    bruc3van/agent-skills-guard/test/test-skills/positive-real/agents-skill-creator Skill

    no license381 repo

    一款提供Agent Skills安全扫描和可视化管理的桌面应用 | A desktop application that provides security scanning and visual management for Agent Skills.

  • Hooks eval

    athola/claude-night-market/plugins/abstract/skills/hooks-eval Skill

    327 repo

    23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context optimization, research, and multi-LLM delegation. 186 skills, 128 commands, 54 agents.

  • Rules eval

    athola/claude-night-market/plugins/abstract/skills/rules-eval Skill

    327 repo

    23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context optimization, research, and multi-LLM delegation. 186 skills, 128 commands, 54 agents.

  • Skills eval

    athola/claude-night-market/plugins/abstract/skills/skills-eval Skill

    327 repo

    23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context optimization, research, and multi-LLM delegation. 186 skills, 128 commands, 54 agents.

  • Skill creator

    syahiidkamil/Software-Engineer-AI-Agent-Atlas/.claude/skills/skill-creator Skill

    no license319 repo

    ATLAS: a senior-engineer layer for Claude Code. Explore with wireframes & prototypes, clarify the essentials, capture it in HTML spec doc then let Claude Code's native plan/goal/workflow loop build. Fewer tokens, less ceremony, faster to what people pictured. KISS/YAGNI/DRY, context decides. No overengineering. Clean architecture that works.

  • Agentsop domain eval set

    agentsope/SkillAlchemy/skills/agentsop-domain-eval-set Skill

    303 repo

    From thought to skill. From signal to structure.

  • Agentsop metric design

    agentsope/SkillAlchemy/skills/agentsop-metric-design Skill

    303 repo

    From thought to skill. From signal to structure.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/ja-JP/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/ja-JP/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/ko-KR/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/ko-KR/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/pt-BR/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/pt-BR/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/tr/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/tr/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/zh-CN/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/zh-CN/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/docs/zh-TW/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Eval harness

    loulanyue/awesome-claude-notes/docs/zh-TW/skills/eval-harness Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.

  • Agent eval

    loulanyue/awesome-claude-notes/skills/agent-eval Skill

    264 repo

    Community-maintained distribution of reusable AI coding agents, commands, skills, hooks, and cross-harness workflows.