Shelves Context & AI engineering
Evals & benchmarks
944 rows from 487 repositories
Deciding whether the output was actually any good, repeatably.
What Evals & benchmarks skills agree on
508 skill files read, by 499 of the 513 authors on this shelf whose files we hold, 2026-09-06
The middle one of the 63 measured here is ~1.3k tokens long, counted with cl100k_base
Counted by distinct author, so one author publishing three of these counts once. Where a claim sits in fewer files than authors, that is said: a claim held by forty authors across three files is one file people copied, not forty people who agreed. Near-identical wordings are grouped and the other wordings are shown, so the grouping is yours to check.
What they tell the agent to do
- Spawn with-skill and baseline runs in the same turn31 of 499 in 24 filesalso worded as Spawn with-skill and baseline runs simultaneously; Launch with-skill and baseline runs together in one turn
- Keep SKILL.md under 500 lines31 of 499 in 24 filesalso worded as Keep the SKILL.md body under 500 lines; Keep SKILL.md under 600 lines using references
- Draft assertions while test runs are in progress31 of 499 in 24 filesalso worded as Draft assertions while runs are in progress; Draft assertions while the runs execute
- Compare against the baseline after changes31 of 499 in 13 filesalso worded as compare current results against a baseline; compare results against the saved baseline
- Define evals before coding26 of 499 in 17 filesalso worded as Define evals before writing code; Define eval criteria before writing code
- Run evals frequently during development25 of 499 in 16 filesalso worded as Run evals frequently; Run evals continuously during development
- Keep evals fast24 of 499 in 15 files
- Version evals with code24 of 499 in 15 filesalso worded as Version evals with the code; Version-control evals alongside code
- Generate the eval viewer before evaluating outputs yourself24 of 499 in 17 filesalso worded as Launch the eval viewer before reviewing outputs yourself; Generate the eval viewer before reviewing outputs yourself
- Generate an eval report after runs24 of 499 in 15 filesalso worded as Generate an eval report after running; Generate a full eval report after runs
- Track pass@k metrics over time22 of 499 in 14 filesalso worded as Track pass@k over time; Track pass@k results over time
- Save a baseline before making changes21 of 499 in 9 filesalso worded as save first-run results as the baseline; Capture a baseline before making changes
- Flag security checks for human review19 of 499 in 10 filesalso worded as Require human review for security checks; Flag security-relevant changes for human review
- Track regressions with each change18 of 499 in 9 filesalso worded as Track regressions with every change
- Measure Core Web Vitals on each target URL18 of 499 in 6 filesalso worded as measure Core Web Vitals on each page
What they tell it not to do
- Do not use other testing skills29 of 499 in 22 filesalso worded as Do not use dedicated testing skills; Don't use other testing skills like /skill-test
- Do not write custom HTML for the viewer26 of 499 in 19 filesalso worded as Don't write custom HTML for the viewer; Do not write custom HTML for the eval viewer
- Never fully automate security checks25 of 499 in 16 filesalso worded as Do not fully automate security checks
- Do not create misleading or malicious skills20 of 499 in 13 filesalso worded as Don't create misleading or malicious skills; Never create misleading or malicious skills
- Do not create workspace directories upfront19 of 499 in 13 filesalso worded as Don't create all workspace directories upfront; Do not create all workspace directories upfront
- Do not force assertions onto subjective outputs14 of 499 in 8 filesalso worded as Don't force assertions onto subjective outputs; Never write subjective assertions
- Do not make negative eval queries obviously irrelevant13 of 499 in 6 filesalso worded as Don't make negative trigger queries obviously irrelevant; Don't make negative eval queries obviously irrelevant
- Do not create skills containing malware or exploit code12 of 499also worded as Do not include malware or exploit code in skills; Never include malware or exploit code in skills
- Do not write assertions when first saving test cases11 of 499 in 10 filesalso worded as Don't draft assertions before running tests; Don't write assertions when first saving test prompts
- Do not edit model cards or model-index10 of 499 in 4 files
What they expect to be installed
- grep51 of 499 in 39 files
- subagents31 of 499 in 24 files
- python330 of 499
- git29 of 499 in 28 files
- generate_review.py28 of 499 in 21 files
- npm run build27 of 499 in 16 files
- npm test25 of 499 in 16 files
- pytest22 of 499 in 19 files
- uv21 of 499 in 13 files
- /eval define21 of 499 in 12 files
What they ask it to produce
- grading.json per run38 of 499 in 31 filesalso worded as grading.json; grading.json per eval
- benchmark.json and benchmark.md35 of 499 in 28 filesalso worded as benchmark.json; benchmark.md
- timing.json per run34 of 499 in 27 filesalso worded as timing.json; timing.json with run statistics
- Eval report33 of 499 in 25 filesalso worded as Full eval report; eval run report
- eval_metadata.json per test case33 of 499 in 26 filesalso worded as eval_metadata.json; evals.json test case file
- Packaged .skill file31 of 499 in 24 filesalso worded as .skill package; .skill package file
- SKILL.md31 of 499 in 24 filesalso worded as SKILL.md draft; Draft SKILL.md
- Eval definition file26 of 499 in 17 filesalso worded as Eval definition markdown file; eval specification file
When Evals & benchmarks authors say to reach for one
The situations these authors wrote into their own files, counted out of the same 499 authors, with the skills that name each one
- User wants to edit or optimize an existing skill38 of 499 in 31 files
- User wants to run evals or benchmark a skill37 of 499 in 30 files
- User wants to create a skill from scratch35 of 499 in 28 files
- User wants to optimize a skill's description for triggering25 of 499 in 24 files
- Skill creator
- and 3 more on this shelf
- Creating regression suites for prompt or agent changes18 of 499 in 10 files
- Eval harness
- Ecc eval harness
- and 11 more on this shelf
- Benchmarking agent performance across model versions18 of 499 in 10 files
- Eval harness
- Ecc eval harness
- and 11 more on this shelf
- Defining pass/fail criteria for task completion18 of 499 in 10 files
- Eval harness
- Ecc eval harness
- and 11 more on this shelf
- When setting up performance baselines18 of 499 in 6 files
- Benchmark
- and 4 more on this shelf
How Evals & benchmarks skills are built
902 skill directories by 448 authors, read from their repositories’ own file trees 2026-08-05
The middle bundle among those shipping files is 5 files, 30.9 KB beside SKILL.md
Counted by distinct author, same as above, so one author publishing forty template copies counts once. SKILL.md itself is not counted as a file, so a single-file skill is one where that file is the whole skill.
The shape
- SKILL.md is the whole skill143 of 448 authors, 400 of 902 skills
- files ship beside it305 of 448 authors, 502 of 902 skills
- executable scripts ship inside199 of 448 authors, 270 of 902 skills
The folders they converge on
- references/226 of 448 authors, 351 of 902 skills
- scripts/178 of 448 authors, 246 of 902 skills
- agents/127 of 448 authors, 178 of 902 skills
- assets/121 of 448 authors, 186 of 902 skills
- eval-viewer/100 of 448 authors, 126 of 902 skills
- evals/34 of 448 authors, 67 of 902 skills
anthropics/skills/skills/skill-creator Skill
in 1 stackno license168,934★ repoPublic repository for Agent Skills
samber/cc-skills-golang/skills/golang-benchmark Skill
2,949★ repo🧑🎨 A collection of Golang agentic skills that works
WithWoz/wozcode-plugin/codex/wozcode/skills/woz-benchmark Skill
no license201★ repoWOZCODE plugin for Claude Code
danielmeppiel/genesis/dev/skills/genesis-evals Skill
61★ repoMarkdown that steers an LLM is code. Genesis is the architectural layer for designing multi-agent, multi-skill systems -- with named patterns, contracts, and substrate portability, before you write them.
liqiongyu/lenny_skills_plus/samples/writing-prds-executable Skill
51★ repo86 agent-executable skill packs converted from RefoundAI’s Lenny skills (unofficial). Works with Codex + Claude Code.
tonyblu331/research-proof/plugins/research-proof-plugin/skills/research-proof Skill
45★ repoPressure-test research claims with falsifiable evidence plans, adversarial checks, frozen verifiers, and proof ledgers.
ContextJet-ai/awesome-llm-observability/skills/add-llm-evals Skill
no license29★ repo50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM apps. Tracing, evals, guardrails, LLMOps.
inngest/inngest-skills/skills/inngest-agent-evals Skill
no license28★ repoAgent Skills for building with Inngest
panaversity/agentfactory-business-plugins/.claude/skills/skill-creator Skill
28★ repoMarketplace of domain-specific plugins for AI agents (Cowork, Claude Code, OpenClaw). Build autonomous business workflows for finance, banking, legal operations, and sales using modular agent skills and commands.
a-tokyo/agent-skills/.agents/skills/skill-creator Skill
15★ repo🧠 AI Agent skills for LLMs and AI Agents - Claude, Codex, Cursor etc.
a-ariff/ariff-claude-plugins/plugins/anti-hallucination/skills/anti-hallucination Skill
14★ repo65 plugins that turn Claude Code into an autonomous development team. 24 agents, 34 skills, 5 hooks. Includes 12-plugin anti-hallucination suite. One-line install.
nagstler/rails-llm-integration Skill
13★🔥🔥 A Claude Skill that teaches Claude Code how to write LLM features
NoesisVision/nasde-toolkit/.claude/skills/nasde-benchmark-calibration Skill
11★ repoCLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini individual subscriptions or API keys.
Tencent-RTC/agent-skills/.claude/skills/trtc-eval Skill
11★ repoTRTC AI Integration Assistant - helps developers integrate Tencent Real-Time Communication SDKs (Chat, Call, RTC Engine, Live, Conference) across Web, Android, iOS, Flutter, and Electron.
di37/EvalSurfer/skills/eval-surfer Skill
11★ repoSkill-first, agent-native evaluation protocol for AI apps
isomoes/skills/.agents/skills/skill-creator Skill
10★ repoCustom skills for agents to support academic paper writing and thesis management
rohit-yadav34/second-mind/.agents/skills/skill-creator Skill
no license8★ repoSecond Mind is a skill that scans your repository and conversation history to capture design decisions, open questions, and session context, storing everything as plain Markdown so project knowledge is portable, human-readable, and easy to resume.
Momo2323-ui/claude-orchestra/skills/hallucination-guard Skill
8★ repoAn operating system for your Claude Code skills, agents & MCPs — organize everything into auto-routing orchestras.
Victory-7291/project-scaffold-setup-skills/.agents/skills/skill-creator Skill
7★ repoReusable agent skills for scaffolding C++ and embedded projects
rusel95/ios-agent-skills/scripts/benchmarking Skill
7★ repoProduction-tested iOS Agent Skills for Claude Code, Codex, and 40+ AI coding tools. 8 enterprise-grade skills covering SwiftUI MVVM, UIKit MVVM, VIPER, TCA, Swift Concurrency, GCD, Testing, and Security Audit.
yiouli/pixie-qa/skills/eval-driven-dev Skill
7★ repoAgent skill for AI agent development
Poorgramer-Zack/dart-expert-skills/.agents/skills/skill-creator Skill
no license7★ repoA comprehensive library of modular Agent Skills for Flutter & Dart development
pitlane-ai/pitlane/skills/testing-with-pitlane Skill
no license6★ repoRace a baseline vs a skill or MCP on real tasks. Hard checks show if it got better, faster, or cheaper. Numbers, not vibes.
crowdin/skills/.agents/skills/skill-creator Skill
5★ repoAgent Skills to help developers using AI agents with Crowdin
yaniv-golan/skill-creator-plus/skill-creator-plus/skills/skill-creator-plus Skill
4★ repoBased on Anthropic's skill-creator — with bug fixes, working Cowork support, and official best practices baked in. Claude skill for creating, testing, and improving other Claude skills.
BayramAnnakov/eval-coach Skill
4★Agent Skill for Evaluation-Driven Development (EDD) - guide AI evaluation strategies with the 50-40-10 rule
build-with-dhiraj/ai-workflow-framework-portability-kit/Plugins/vercel-marketplace-source/.claude/skills/benchmark-agents Skill
4★ repoPortable, self-contained snapshot of a complete Claude Code setup — 36 specialist agents, 134 skills, plugins, MCP servers & host tooling. Clone, claude login, run one script, restore the whole orchestration stack in ~20 min.
stempeck/agentfactory/.claude/skills/agentic-skill-eval Skill
3★ repoMulti-agent orchestration CLI for Claude Code — declarative TOML workflows, autonomous agents, context-compression recovery, inter-agent mail.
tmuskal/arc-agi-benchmarker/.claude/skills/benchmark-adder Skill
no license3★ repoProve you achieved AGI at home by testing your claude-code setup against arc-agi-3 benchmarks
urmzd/dotfiles/dot_agents/skills/agent-design-doctrine Skill
3★ repoCross-platform dotfiles managed by Chezmoi with Homebrew/apt and per-language version managers. One-command bootstrap for macOS and Linux with Neovim, Tmux, Zsh, and AI agent skills.
drsh4dow/grug-brain-skill/.agents/skills/skill-creator Skill
3★ repoA simplicity-first skill for reviewing architecture, refactors, APIs, and testing strategy without adding unnecessary complexity.
jmcentire/signet-eval/io.github.jmcentire/signet-eval MCP server
3★ repoDeterministic policy enforcement and MCP management for AI agent tool calls.
harshitsinghbhandari/domain-expansion/.agents/skills/skill-creator Skill
no license2★ repoCollection of Claude agent skills (code audits, LLM councils, PR review, resume tooling) installable individually via npx skills add.
abhishekgahlot2/agent-optimization/skills/agent-optimization Skill
2★ repoPrompt, checklist, and Agent Skill for optimizing LLM agents without benchmark hacks.
unshDee/proofrag/skills/proofrag Skill
2★ repoPoint your agent at your docs and your RAG app; get a golden test set + an LLM-as-judge & retrieval scorecard, in one command.
croftspan/gigo/.claude/skills/eval Skill
1★ repo義剛 GIGO. AI projects that get better every session, not worse.
Jaypatel1511/cdfi-superpowers/skills/cdfi-peer-benchmark Skill
1★ repoAI skills for NMTC eligibility, bank-CDFI peer benchmarking & HMDA analysis — grounded in audited PyPI tools, not hallucinated.
Modellix/modellix-skill/.agents/skills/skill-creator Skill
1★ repoAn Agent Skill of Modellix documentation and capabilities reference.
varunk130/AI-Eval-Skills/skills/cost-quality-frontier Skill
1★ repoCurated AI agent evaluation skills from Microsoft's Eval Guide — plan, generate, run, and interpret eval suites for Copilot Studio agents
vindm/dotclaude/deferred/skills/ai-workflow Skill
1★ repoAI dev infrastructure framework for Claude Code. /dotclaude:bootstrap authors CLAUDE.md + docs/ + .claude/ tuned to your project. Per-domain: design (showpiece), coding, planning, testing, data, ai-workflow.
arcjet/skills/.agents/skills/skill-creator Skill
1★ repoSkills to make your AI coding agent an Arcjet security expert.
morzecrew/agent-skills/.agents/skills/skill-creator Skill
1★ repoCollection of agent skills
ahnafyy/skills-evals/skills/setup-skills-evals Skill
1★ repoCatch when your agent skills, instructions, and rules stop triggering — or stop working.
Skill creator anthropics 20260226
MatrixFounder/Universal-skills/archive/skill-creator-anthropics-20260226 Skill
1★ repoCollection of high-leverage "Meta-Skills" designed to upgrade AI Agents from simple chat bots to autonomous engineers
jtmthf/skills/.agents/skills/skill-creator Skill
no license1★ repoPersonal Agent Skills
Metaverse-Cloud/EngageLab-Skills/.agents/skills/skill-creator Skill
1★ repoEngageLab-Skills
zinan92/repo-evals Skill
1★Claim-first repo 评测框架。in target repo + claim map → out bilingual verdict dossier + all-evals dashboard
krkn-s/ai-workspace/.pi/skills/skill-creator Skill
1★ repoKRKN's Personal AI workspace • Skills / Prompts / ...
eagerworks/skills/.agents/skills/skill-creator Skill
1★ repoA collection of portable agent skills by Eagerworks — install via the skills.sh CLI for Claude Code, Cursor, Copilot, Codex, Amp, and more.
0SxD/candidacy-checker-eval Skill
1★Generic two-layer candidacy evaluation framework: Socratic profile-discovery loop with dialectic confidence saturation, plus a flat boolean rubric.
Xtqm/skills/skills/vercel-eve-framework Skill
1★ repoI'll be adding my personal agent skills into this repo.
udohjeremiah/skills/.agents/skills/skill-creator Skill
1★ repoReusable Agent Skills for AI coding agents, focused on software engineering workflows.
Paldom/github-skills/.claude/skills/add-skill Skill
1★ repoAgent Skills for professional GitHub repos - brief, structured, SEO-friendly READMEs that convert visitors into users, plus complete OSS scaffolding: community health files, templates, and discoverability best practices.
Paldom/icon-designer-skills/.claude/skills/add-skill Skill
1★ repoAgent Skills that design minimalist app and OSS package icons from a text brief or project context - symmetric logos on dark grey, Apple-style rounded-rectangle backgrounds.
Paldom/skillskit/skills/add-skill Skill
1★ repoFrom context to installable agent skills - research packs in, validated skills.sh-ready skills out. Scaffolding, eval-first authoring, validation, and deployment included.
Paldom/python-skills/.claude/skills/add-skill Skill
1★ repoAgent Skills and hooks for maintaining high-quality open-source Python packages - an agentic engineering setup covering linting, testing, packaging, releases, and CI quality gates.
Paldom/node-skills/.claude/skills/add-skill Skill
1★ repoAgent Skills for maintaining high-quality open-source Node.js, TypeScript, Next.js, and React apps and packages - linting, type safety, testing, packaging and releases, and CI quality gates.
nathan8823/fairbill/.claude/skills/benchmark-price Skill
1★ repoTurn your AI agent into a medical-bill negotiation advocate — an open-source playbook of skills, letters, and verified patient rights. Pay what's fair, nothing more.
cyanheads/evals-mcp-server/io.github.cyanheads/evals-mcp-server MCP server
1★ repoAuthor verifiable eval records through a draft→review→revise→submit loop with enforced graders.
jpoindexter/reward-seeking-safety-skills/skills/contrastive-authority-eval Skill
0★ repoAgent skills and eval utilities for reward-seeking, grader-targeting, and oversight-dependent behavior.