Evals & benchmarks
946 rows, by stacks then stars
aoli0919/learn-anything-skills/skills/hallucination-checker Skill
20 days old0★ repoBeginner-first agent skills that turn 'I want to learn X' into a 30-day path, tutor loop, projects, and learning memory.
aoli0919/research-radar-skills/skills/benchmark-watch Skill
19 days old0★ repoSkills for tracking fast-moving research without drowning in arXiv, benchmarks, and lab updates.
aoli0919/verification-trust-skills/skills/hallucination-checker Skill
19 days old0★ repoSkills for checking claims, auditing sources, detecting hallucinations, and labeling uncertainty.
ats4321/claude-engineering-skills/skills/evaluation-frameworks Skill
0★ repo26 repository-agnostic engineering skills for Claude Code — debugging, design, review, validation, and AI engineering as operational runbooks.
avnath13/evalpilot/skills/eval-audit Skill
20 days old0★ repoAgent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.
avnath13/evalpilot Skill
20 days old0★Agent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.
avnath13/evalpilot/skills/grade Skill
20 days old0★ repoAgent evals on autopilot: find quality bugs in your AI agent, ship a targeted fix, and prove it on a held-out set. Zero-dependency Agent Skill + CLI.
awesome-liuxiao/agent-skill-doctor/benchmarks/public/v1/fixtures/dynamic-eval Skill
17 days old0★ repoFind broken and risky Codex or Claude Code skills locally
awesome-liuxiao/agent-skill-doctor/benchmarks/public/v1/upstream/anthropic/skill-creator Skill
17 days old0★ repoFind broken and risky Codex or Claude Code skills locally
bitrails-dev/skills/gsd/skills/gsd-eval-review Skill
no license0★ repoHandy AI skills
blakebauman/skillist-validator/skills/skill-description-tuning Skill
12 days old0★ repoValidate AI coding assistant skills against the Agent Skills specification (agentskills.io) — a dependency-free validator with fix hints, plus skills for auditing instruction quality and tuning description triggering.
cfregly/gpu-perf-tune/plugins/profile-and-optimize/skills/inference-model-eval Skill
0★ repo31 GPU inference profiling and optimization skills for Claude Code, with a bundled MCP server
chienchuanw/chuan-skills/plugins/skill-optimize/skills/skill-benchmark Skill
no license0★ repoPersonal plugin marketplace of Claude Code skills (slash commands)
chobizzy/llm-wiki/skills/skill-creator Skill
22 days old0★ repoAn LLM wiki for Obsidian: a markdown knowledge base your AI agents compile and maintain under written law. 32 skills for Claude Code and Hermes, a zero-dependency Python CLI, and a vault template governed by an 11-law constitution.
claudette-agent/agent-eval-kit Skill
no license0★Free mini-eval and scorecard for AI coding agents: Claude Code, Codex, Cursor, Copilot, Windsurf.
clauxel/gemini-upgrade-qa-mcp/com.clauxel.geminiupgradeqa/geminiupgradeqa-mcp MCP server
0★ repoRemote MCP for Gemini upgrade evals, prompt regressions, output diffs, and eval receipts.
cleatai/agent-skills/skills/price-benchmark Skill
24 days old0★ repoInstallable agent skills (SKILL.md) for government contracting. Claude Code plugin + copyable playbooks driving the CLEATUS MCP server.
cmdecker95/skills/skills/skill-creator Skill
no license0★ repo🦾 My own little curation of skills
cody-hutson/pmo-platform/core/skills/eval-writer Skill
no license0★ repoA modular PMO & release-management platform for Claude Code: skills, governance disciplines, and a 13-stage release pipeline.
CohenD/cc-stack/.claude/skills/promptfoo-evals Skill
no license0★ repoPre-wired Claude Code workspace: Next.js 16 + AI SDK 7 + shadcn/ui + Tailwind v4 + DuckDB, with Skills & MCP servers ready on clone.
cskwork/skill-ab-eval Skill
0★Prove whether a SKILL.md actually changes agent behavior (with/without) and which CLI harness does a task best — agent-native subagents or CLI. agentskills.io-compatible. Site: https://cskwork.github.io/skill-ab-eval/
Curtisflo/karyon/skills/benchmark-leakage Skill
no license0★ repoLegible, deterministic QC/qualification for bio-AI tool outputs — named-reason contracts + agent skills that compose with NVIDIA BioNeMo.
Cy4nLiang/claude-code-prompt-architect/skills/pa-eval Skill
0★ repoPrompt optimizer & compiler for Claude Code — intent mining, genre-aware skeletons, multi-candidate + LLM-judge eval. Text / image / video prompts (Midjourney, Seedance, Sora, Kling). 新手友好的提示词优化套件
dbsectrainer/mcp-eval-runner/io.github.dbsectrainer/mcp-eval-runner MCP server
0★ repoA standardized testing harness for MCP servers and agent workflows
dills122/ai-central/templates/skills/imported/claude-skills/engineering/skills/self-eval Skill
0★ repoCentral library for AI coding context: steering files, AGENTS templates, reusable skills, scaffold scripts, and guided setup for new or existing projects.
DrOlu/agent-skills/skills/inngest-agent-evals Skill
no license0★ repoOpen agent skills for the skills.sh ecosystem — browser, docs, mail, media, security, networking, orchestration, and more.
DrOlu/agent-skills/skills/inngest-brownfield-audit Skill
no license0★ repoOpen agent skills for the skills.sh ecosystem — browser, docs, mail, media, security, networking, orchestration, and more.
Drvivek34/Skill-Bazaar/coding-development/skill-creator Skill
no license0★ repo🛒 Skill Bazaar — an all-in-one open collection of AI Agent Skills (SKILL.md), organized by category. Part of the Mega AI Bazaar.
ekreloff/claude-self-eval/skills/self-eval Skill
no license0★ repoHonest AI work evaluation for Claude Code — two-axis scoring with anti-inflation mechanisms
epitech-toulouse/epitech-pre-eval Skill
no license0★Source command megaminx eval v
Erlemar/cayley-puzzles/.agents/skills/source-command-megaminx-eval-v Skill
no license0★ repoNeural distance heuristics + GPU/TPU beam search for the CayleyPy IHES Picture Cube and Megaminx puzzles
Evan-Daruwalla/claude-skill-suite/llm-eval-harness Skill
27 days old0★ repoClaude Code skills for running models cost-effectively: security gates (secret scanner, commit-gate), model-quality tooling (eval harness, token-squeeze, compact-io, opus-workers), review/advisory (trusted-advisor, audit, skill-vet, research-brief), and a read-only reorg-proposal advisor.
fabioc-aloha/Alex_ACT_Edition/.github/skills/anti-hallucination Skill
0★ repoACT-Edition brain template for AI coding assistants — critical thinking, epistemic calibration, and structured reasoning
fbsmna-coder/karpathy-pro-max/skills/anti-hallucination Skill
0★ repoStop Claude Code from hallucinating — Karpathy-grade discipline in 8 skills
fengluisobel/ask-dongfeng-hermes/skills/mlops/evaluation/lm-evaluation-harness Skill
0★ repoHermes Agent distribution with Ask DongFeng bundled as a built-in control-loop framework skill
fermeridamagni/skills/.agents/skills/skill-creator Skill
0★ repoA collection of AI skills created for me
Frog1205/OPC-Skills/skills/req-eval Skill
0★ repoOPC-Skills — AI 一人公司(One-Person Company)Claude Code 技能合集。一个人 + AI = 一支团队。战略/法务/内容/文档等虚拟部门,18 个开箱即用的 AI 技能。
getappniche/aso-skills/skills/revenue-benchmark Skill
12 days old0★ repoASO & app-market research skills for AI agents — Claude Code, Cursor, and any MCP client. Install: npx skills add getappniche/aso-skills
goharabbas321/zeoel/skills/agent-eval Skill
no license0★ repoZeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.
goharabbas321/zeoel/skills/benchmark Skill
no license0★ repoZeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.
goharabbas321/zeoel/skills/eval-harness Skill
no license0★ repoZeoel-AI is a production-grade, multi-agent SaaS software engineering agency. Instead of single, fragile LLM prompts that lose context, Zeoel-AI coordinates a complete team of 33 specialized agents (UX designers, database architects, systems engineers, security auditors, QA engineers, etc.) to brainstorm, plan, build, and verify entire software.
goharabbas321/zeoel-framework/all-skills/agent-eval Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/all-skills/awesome-claude-skills/composio-skills/benchmark-email-automation Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/all-skills/benchmark Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/all-skills/composio-skills/benchmark-email-automation Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/all-skills/eval-harness Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/all-skills/healthcare-eval-harness Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/agent-eval Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/awesome-claude-skills/composio-skills/benchmark-email-automation Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/benchmark Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/composio-skills/benchmark-email-automation Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/eval-harness Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
goharabbas321/zeoel-framework/.agents/skills/zeoel/skills/healthcare-eval-harness Skill
no license0★ repo23-agent AI development team for Claude Code, Cursor, and Gemini CLI. Multi-agent orchestration framework with strict TDD, sprint planning, 420+ skills, and automated QA/security/SEO audits. Stop vibe-coding — start shipping production-grade software.
GRIDLOCK-NYC/claude-skills/skills/skill-builder Skill
0★ repo22 production-tested Claude Code skills: code review, planning, session audits, skill builders, and more.
GRIDLOCK-NYC/claude-skills/skills/skill-eval Skill
0★ repo22 production-tested Claude Code skills: code review, planning, session audits, skill builders, and more.
g-shevchenko/agentic-quality-skills/skills/golden-benchmark-uplift-loop Skill
0★ repoProduction-grade quality skills for AI coding agents: red-first TDD, quality gates, and golden benchmark uplift loops.
guynhsichngeodiec/cc-skills-golang/skills/golang-benchmark Skill
0★ repoExtend Go coding assistants with reusable skills for language, testing, security, and observability in production-ready Golang projects
h3y6e/agent-skills/.vendor/skills/agent-platform-eval-flywheel Skill
0★ repomy agent skills
h3y6e/agent-skills/.vendor/skills/system-benchmark Skill
0★ repomy agent skills
h3y6e/agent-skills/.vendor/skills/waxa-eval Skill
0★ repomy agent skills