agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Hallucination checker

    aoli0919/learn-anything-skills/skills/hallucination-checker Skill

    20 days old0 repo

    Beginner-first agent skills that turn 'I want to learn X' into a 30-day path, tutor loop, projects, and learning memory.

  • Benchmark watch

    aoli0919/research-radar-skills/skills/benchmark-watch Skill

    19 days old0 repo

    Skills for tracking fast-moving research without drowning in arXiv, benchmarks, and lab updates.

  • Hallucination checker

    aoli0919/verification-trust-skills/skills/hallucination-checker Skill

    19 days old0 repo

    Skills for checking claims, auditing sources, detecting hallucinations, and labeling uncertainty.

  • Evaluation frameworks

    ats4321/claude-engineering-skills/skills/evaluation-frameworks Skill

    0 repo

    26 repository-agnostic engineering skills for Claude Code — debugging, design, review, validation, and AI engineering as operational runbooks.

  • Eval audit

    avnath13/evalpilot/skills/eval-audit Skill

    20 days old0 repo

    Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.

  • Evalpilot

    avnath13/evalpilot Skill

    20 days old0

    Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.

  • Grade

    avnath13/evalpilot/skills/grade Skill

    20 days old0 repo

    Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.

  • Dynamic eval

    awesome-liuxiao/agent-skill-doctor/benchmarks/public/v1/fixtures/dynamic-eval Skill

    17 days old0 repo

    Find broken and risky Codex or Claude Code skills locally

  • Skill creator

    awesome-liuxiao/agent-skill-doctor/benchmarks/public/v1/upstream/anthropic/skill-creator Skill

    17 days old0 repo

    Find broken and risky Codex or Claude Code skills locally

  • Gsd eval review

    bitrails-dev/skills/gsd/skills/gsd-eval-review Skill

    no license0 repo

    Handy AI skills

  • Skill description tuning

    blakebauman/skillist-validator/skills/skill-description-tuning Skill

    12 days old0 repo

    Validate AI coding assistant skills against the Agent Skills specification (agentskills.io) — a dependency-free validator with fix hints, plus skills for auditing instruction quality and tuning description triggering.

  • Inference model eval

    cfregly/gpu-perf-tune/plugins/profile-and-optimize/skills/inference-model-eval Skill

    0 repo

    31 GPU inference profiling and optimization skills for Claude Code, with a bundled MCP server

  • Skill benchmark

    chienchuanw/chuan-skills/plugins/skill-optimize/skills/skill-benchmark Skill

    no license0 repo

    Personal plugin marketplace of Claude Code skills (slash commands)

  • Skill creator

    chobizzy/llm-wiki/skills/skill-creator Skill

    22 days old0 repo

    An LLM wiki for Obsidian: a markdown knowledge base your AI agents compile and maintain under written law. 32 skills for Claude Code and Hermes, a zero-dependency Python CLI, and a vault template governed by an 11-law constitution.

  • Agent eval kit

    claudette-agent/agent-eval-kit Skill

    no license0

    Free mini-eval and scorecard for AI coding agents: Claude Code, Codex, Cursor, Copilot, Windsurf.

  • Geminiupgradeqa mcp

    clauxel/gemini-upgrade-qa-mcp/com.clauxel.geminiupgradeqa/geminiupgradeqa-mcp MCP server

    0 repo

    Remote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.

  • Price benchmark

    cleatai/agent-skills/skills/price-benchmark Skill

    24 days old0 repo

    Installable agent skills (SKILL.md) for government contracting. Claude Code plugin + copyable playbooks driving the CLEATUS MCP server.

  • Skill creator

    cmdecker95/skills/skills/skill-creator Skill

    no license0 repo

    🦾 My own little curation of skills

  • Eval writer

    cody-hutson/pmo-platform/core/skills/eval-writer Skill

    no license0 repo

    A modular PMO & release-management platform for Claude Code: skills, governance disciplines, and a 13-stage release pipeline.

  • Promptfoo evals

    CohenD/cc-stack/.claude/skills/promptfoo-evals Skill

    no license0 repo

    Pre-wired Claude Code workspace: Next.js 16 + AI SDK 7 + shadcn/ui + Tailwind v4 + DuckDB, with Skills & MCP servers ready on clone.

  • Skill ab eval

    cskwork/skill-ab-eval Skill

    0

    Prove whether a SKILL.md actually changes agent behavior (with/without) and which CLI harness does a task best — agent-native subagents or CLI. agentskills.io-compatible. Site: https://cskwork.github.io/skill-ab-eval/

  • Benchmark leakage

    Curtisflo/karyon/skills/benchmark-leakage Skill

    no license0 repo

    Legible, deterministic QC/qualification for bio-AI tool outputs — named-reason contracts + agent skills that compose with NVIDIA BioNeMo.

  • Pa eval

    Cy4nLiang/claude-code-prompt-architect/skills/pa-eval Skill

    0 repo

    Prompt optimizer & compiler for Claude Code — intent mining, genre-aware skeletons, multi-candidate + LLM-judge eval. Text / image / video prompts (Midjourney, Seedance, Sora, Kling). 新手友好的提示词优化套件

  • Mcp eval runner

    dbsectrainer/mcp-eval-runner/io.github.dbsectrainer/mcp-eval-runner MCP server

    0 repo

    A standardized testing harness for MCP servers and agent workflows

  • Self eval

    dills122/ai-central/templates/skills/imported/claude-skills/engineering/skills/self-eval Skill

    0 repo

    Central library for AI coding context: steering files, AGENTS templates, reusable skills, scaffold scripts, and guided setup for new or existing projects.

  • Inngest agent evals

    DrOlu/agent-skills/skills/inngest-agent-evals Skill

    no license0 repo

    Open agent skills for the skills.sh ecosystem — browser, docs, mail, media, security, networking, orchestration, and more.

  • Inngest brownfield audit

    DrOlu/agent-skills/skills/inngest-brownfield-audit Skill

    no license0 repo

    Open agent skills for the skills.sh ecosystem — browser, docs, mail, media, security, networking, orchestration, and more.

  • Skill creator

    Drvivek34/Skill-Bazaar/coding-development/skill-creator Skill

    no license0 repo

    🛒 Skill Bazaar — an all-in-one open collection of AI Agent Skills (SKILL.md), organized by category. Part of the Mega AI Bazaar.

  • Self eval

    ekreloff/claude-self-eval/skills/self-eval Skill

    no license0 repo

    Honest AI work evaluation for Claude Code — two-axis scoring with anti-inflation mechanisms

  • Epitech pre eval

    epitech-toulouse/epitech-pre-eval Skill

    no license0
  • Source command megaminx eval v

    Erlemar/cayley-puzzles/.agents/skills/source-command-megaminx-eval-v Skill

    no license0 repo

    Neural distance heuristics + GPU/TPU beam search for the CayleyPy IHES Picture Cube and Megaminx puzzles

  • Llm eval harness

    Evan-Daruwalla/claude-skill-suite/llm-eval-harness Skill

    27 days old0 repo

    Claude Code skills for running models cost-effectively: security gates (secret scanner, commit-gate), model-quality tooling (eval harness, token-squeeze, compact-io, opus-workers), review/advisory (trusted-advisor, audit, skill-vet, research-brief), and a read-only reorg-proposal advisor.

  • Anti hallucination

    fabioc-aloha/Alex_ACT_Edition/.github/skills/anti-hallucination Skill

    0 repo

    ACT-Edition brain template for AI coding assistants — critical thinking, epistemic calibration, and structured reasoning

  • Anti hallucination

    fbsmna-coder/karpathy-pro-max/skills/anti-hallucination Skill

    0 repo

    Stop Claude Code from hallucinating — Karpathy-grade discipline in 8 skills

  • Lm evaluation harness

    fengluisobel/ask-dongfeng-hermes/skills/mlops/evaluation/lm-evaluation-harness Skill

    0 repo

    Hermes Agent distribution with Ask DongFeng bundled as a built-in control-loop framework skill

  • Skill creator

    fermeridamagni/skills/.agents/skills/skill-creator Skill

    0 repo

    A collection of AI skills created for me

  • Req eval

    Frog1205/OPC-Skills/skills/req-eval Skill

    0 repo

    OPC-Skills — AI 一人公司(One-Person Company)Claude Code 技能合集。一个人 + AI = 一支团队。战略/法务/内容/文档等虚拟部门,18 个开箱即用的 AI 技能。

  • Revenue benchmark

    getappniche/aso-skills/skills/revenue-benchmark Skill

    12 days old0 repo

    ASO & app-market research skills for AI agents — Claude Code, Cursor, and any MCP client. Install: npx skills add getappniche/aso-skills

  • Agent eval

    goharabbas321/zeoel/skills/agent-eval Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Benchmark

    goharabbas321/zeoel/skills/benchmark Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Eval harness

    goharabbas321/zeoel/skills/eval-harness Skill

    no license0 repo

    Zeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.

  • Agent eval

    goharabbas321/zeoel-framework/all-skills/agent-eval Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark email automation

    goharabbas321/zeoel-framework/all-skills/awesome-claude-skills/composio-skills/benchmark-email-automation Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark

    goharabbas321/zeoel-framework/all-skills/benchmark Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark email automation

    goharabbas321/zeoel-framework/all-skills/composio-skills/benchmark-email-automation Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Eval harness

    goharabbas321/zeoel-framework/all-skills/eval-harness Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Healthcare eval harness

    goharabbas321/zeoel-framework/all-skills/healthcare-eval-harness Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Agent eval

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/agent-eval Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark email automation

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/awesome-claude-skills/composio-skills/benchmark-email-automation Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/benchmark Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Benchmark email automation

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/composio-skills/benchmark-email-automation Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Eval harness

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/eval-harness Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Healthcare eval harness

    goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/healthcare-eval-harness Skill

    no license0 repo

    23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.

  • Skill builder

    GRIDLOCK-NYC/claude-skills/skills/skill-builder Skill

    0 repo

    22 production-tested Claude Code skills: code review, planning, session audits, skill builders, and more.

  • Skill eval

    GRIDLOCK-NYC/claude-skills/skills/skill-eval Skill

    0 repo

    22 production-tested Claude Code skills: code review, planning, session audits, skill builders, and more.

  • Golden benchmark uplift loop

    g-shevchenko/agentic-quality-skills/skills/golden-benchmark-uplift-loop Skill

    0 repo

    Production-grade quality skills for AI coding agents: red-first TDD, quality gates, and golden benchmark uplift loops.

  • Golang benchmark

    guynhsichngeodiec/cc-skills-golang/skills/golang-benchmark Skill

    0 repo

    Extend Go coding assistants with reusable skills for language, testing, security, and observability in production-ready Golang projects

  • Agent platform eval flywheel

    h3y6e/agent-skills/.vendor/skills/agent-platform-eval-flywheel Skill

    0 repo

    my agent skills

  • System benchmark

    h3y6e/agent-skills/.vendor/skills/system-benchmark Skill

    0 repo

    my agent skills

  • Waxa eval

    h3y6e/agent-skills/.vendor/skills/waxa-eval Skill

    0 repo

    my agent skills