Evals & benchmarks
946 rows, by stacks then stars
mj-deving/pai-skills/skills/Utilities Skill
no license0★ repoCurated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).
mj-deving/pai-skills/skills/Utilities/Evals Skill
no license0★ repoCurated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).
Moliboy5000/.claude/skills/benchmark Skill
no license0★ repoClaude dotfile containing skills, agents etc I use.
Moliboy5000/.claude/skills/benchmark-models Skill
no license0★ repoClaude dotfile containing skills, agents etc I use.
Moliboy5000/.claude/skills/skill-creator Skill
no license0★ repoClaude dotfile containing skills, agents etc I use.
morzecrew/agent-skills/.agents/skills/skill-creator Skill
0★ repoCollection of agent skills
mustafayilmazart/kesif-claude-skills/skills/skill-creator Skill
0★ repo14 production-ready Claude Code skill — TTS, video, NotebookLM, YouTube, token optimization, prompt engineering
NeuralMedic-DE/claude-skills/rag-eval-guardrails Skill
0★ repoVerified-gate Claude Code skills that run their checks and prove the result — by NeuralMedic. Web, accessibility, healthcare & compliance, data, IaC, Odoo.
Nyota-tree/pm-trace-eval/skills/pm-trace-eval Skill
25 days old0★ repoPM-facing AI conversation trace eval skill: north-star metrics, 1-5 scores, red/yellow/green, per-turn quote review
onyx679/automotive-ve-ai-skills-kit/skills/ve-skill-benchmark-adapter Skill
0★ repoAI Skill templates and workflow scoring tools for automotive value engineering and VAVE productivity scenarios
pagoda111king/claude-code-pack/template/.claude/skills/eval-refine Skill
no license0★ repoClaude Code config pack · 11 agents · 10 skills · 10 commands · 4 hooks · CLAUDE.md template · Solo Founder OS
pagoda111king/claude-code-pack/template/.claude/skills/idea-eval Skill
no license0★ repoClaude Code config pack · 11 agents · 10 skills · 10 commands · 4 hooks · CLAUDE.md template · Solo Founder OS
Paldom/landing-page-builder-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for building high-converting, modern, performant landing pages: design tokens, conversion copywriting, marketing UX, layout patterns, Next.js/React + shadcn implementation, scroll motion, and performance/SEO/accessibility.
Paldom/llm-council-skills/.claude/skills/add-skill Skill
0★ repoAgent Skills for building and operating LLM councils - multi-model deliberation with anonymized peer review, robust aggregation, calibrated escalation, and defenses against correlated failure.
Paldom/mcp-server-skills/.claude/skills/add-skill Skill
17 days old0★ repoAgent Skills for designing and implementing Model Context Protocol (MCP) servers - server architecture, tools, resources, prompts, transports, authorization, testing, and deployment.
Paldom/noslop/.claude/skills/add-skill Skill
17 days old0★ repoAgent Skills that strip AI writing tells from copy - em dashes, 'it's not X, it's Y' constructions, hedging, and bloat - so text reads like a human wrote it, with evals to verify.
Paldom/promptimize/.claude/skills/add-skill Skill
18 days old0★ repoAgent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.
Paldom/promptimize/skills/evaluate-prompts Skill
18 days old0★ repoAgent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.
Paldom/screenshooter/.claude/skills/add-skill Skill
16 days old0★ repoSkills for creating polished screen-recording walkthrough and tutorial videos of web apps, with smooth scripted mouse movements, zoom effects, and explanation cards.
Paldom/stealth-browser-skills/.claude/skills/add-skill Skill
17 days old0★ repoAgent Skills for stealthy, human-like browser automation with Playwright and Camoufox — real-browser agentic browsing that evades bot detection.
Paldom/terminaltor/.claude/skills/add-skill Skill
15 days old0★ repoAgent Skills for recording, sanitizing, and rendering terminal session demos - capture real command walkthroughs with agents, redact sensitive output, and render polished casts for READMEs and docs.
Pdbjork/hermosskills-site/skills/ecc-eval-harness Skill
29 days old0★ repoHermesskills — a curated marketplace & repository of AI-agent skills. OpenClaw-compatible catalog at /catalog/v1/skills.json. 30% of commissions fund public-good AI.
Evaluation lm evaluation harness
percymcn/agent-cookbook/skills/mlops/evaluation-lm-evaluation-harness Skill
0★ repoProduction-ready AI agent skills, playbooks, and workflows from the pharma6 automation lab
pgoell/pgoell-claude-tools/plugins/agent-system-management/skills/creating-skills Skill
0★ repoCollection of my personal claude skills
planifest/planifest-framework/planifest-framework/external-skills/benchmark-models Skill
0★ repoA specification framework for agentic development. Agents build from complete specs - not guesses.
planifest/planifest-framework/planifest-framework/external-skills/benchmark Skill
0★ repoA specification framework for agentic development. Agents build from complete specs - not guesses.
PlevanTem/luban-skill/.claude/skills/ai-pm Skill
0★ repo不是泛泛而谈,蒸馏一个专家,带走ta的方法论和专业判断。Distill expert methodology into agent kits — not vibes persona prompts
ploteddie-bit/skills/skill-creator Skill
no license0★ repoCollection de compétences agentiques modulables pour Kimi, Aegis et assistants IA
pngdeity/apm-user-repository/packages/skill-eval-pipeline/.apm/skills/skill-eval-pipeline Skill
0★ repoPersonal APM marketplace — skills, prompts, agents, and instructions for AI coding agents
po4yka/llm-wiki-skills/skills/llm-wiki-eval Skill
0★ repoPortable Agent Skills for building, operating, evaluating, and governing LLM-Wiki knowledge systems.
po4yka/llm-wiki-skills/skills/llm-wiki-eval-tooling Skill
0★ repoPortable Agent Skills for building, operating, evaluating, and governing LLM-Wiki knowledge systems.
robcsaszar/ai-forge/skills/ai-forge-eval Skill
26 days old0★ repoClaude Code skills for creating, judging, evaluating, and updating skills and agents
robdasi/skills/eval-set Skill
0★ repoFree, working Claude skills I use to run an AI automation studio. Drop-in SKILL.md files. By Robin Laires / Laires Labs.
rulebased-io/claude-plugin/packages/harness/skills/eval-log Skill
0★ repoClaude Code plugins by rulebased.io — harness audit/init/recommend + second-brain capture/connect/review/organize
rulebased-io/harness/rulebased-harness/skills/eval-log Skill
0★ repoHarness engineering tools for AI agents — audit, init, recommend your project's agent-readiness (AGENTS.md, context engineering, eval)
SamyakJhaveri/loam/seed/_research/skills/eval-grader Skill
0★ repoCopier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.
SamyakJhaveri/loam/seed/_research/skills/eval-run Skill
0★ repoCopier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.
SamyakJhaveri/loam/seed/_research/skills/overnight-eval Skill
0★ repoCopier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.
SamyakJhaveri/loam/seed/_research/skills/post-eval Skill
0★ repoCopier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.
Sapien3/amortized-intelligence Skill
0★A Claude Code skill: spend frontier-model tokens at design time, compile them into artifacts — harnesses, rubrics, evals, skills — that cheap models run at near-zero marginal cost.
satishTheLegend/ai-rmf-governor Skill
0★Claude Code skill: run NIST AI RMF (Govern/Map/Measure/Manage) on your LLM/ML feature — named failure modes, metric thresholds, an offline eval gate, and signed-off residual risk.
satishTheLegend/eval-harness-architect Skill
0★Stand up a rigorous, regression-proof evaluation harness for any LLM/agent system from zero — datasets, scorers, CI gates, and drift monitoring.
search-atlas-group/amm-founding-circle/skills/instinct-learn-eval Skill
0★ repoThe AMM founding-circle home base: 36 Claude skills (AEO/SEO + agentic engineering + security), the agentic ladder, playbooks, and automations. No paid APIs required.
Sheshiyer/skill-clusters/skills/agent-eval Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/agentic-engineering Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/benchmark Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/eval-harness Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/evals Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/healthcare-eval-harness Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/hugging-face-community-evals Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/optimize Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/quality-eval-core Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
Sheshiyer/skill-clusters/skills/quality-eval-orchestrator Skill
0★ repoHub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.
sisodiabhumca/agent-skills/skills/ai-eval-regression-tester Skill
0★ repoProduction-Ready Agent Skills : product analytics, growth experiments, CRM, research synthesis, postmortems, data contracts, SaaS spend, compliance, architecture maps, and LLM eval and many more.
SkillMedev/ai-engineer-toolkit/skills/llm-evaluation Skill
0★ repoShip LLM features that work — prompts, evals, agents, and MCP servers.
soltraveler-sri/Steerable-Protocol/skills/eval-authoring Skill
29 days old0★ repoAn open standard for making apps steerable, shipped with a reference runtime and a skill that lets a coding agent implement it in your codebase.
songsunny00/MySkills/BaseSkills/.agents/skills/skill-creator Skill
no license0★ repo根据不同角色,收集或编写日常常用skill
sorinLupaIonut/digital-fte/.claude/skills/skill-creator Skill
15 days old0★ repoA customer-support Digital FTE (AI Worker): OpenAI Agents SDK on a sandboxed runtime, capabilities as portable Skills, a scoped customer-data MCP server instead of raw SQL, pgvector semantic search on Neon Postgres, and a transactional audit row behind every action.
srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Business & Ops/cross-eval Skill
0★ repoNative macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.
srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Data & AI/eval Skill
0★ repoNative macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.