agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Utilities

    mj-deving/pai-skills/skills/Utilities Skill

    no license0 repo

    Curated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).

  • Evals

    mj-deving/pai-skills/skills/Utilities/Evals Skill

    no license0 repo

    Curated, sanitized export of 21 agent-skill packages for Claude Code and Codex, gated by an automated publication audit (no secrets, no local paths).

  • Benchmark

    Moliboy5000/.claude/skills/benchmark Skill

    no license0 repo

    Claude dotfile containing skills, agents etc I use.

  • Benchmark models

    Moliboy5000/.claude/skills/benchmark-models Skill

    no license0 repo

    Claude dotfile containing skills, agents etc I use.

  • Skill creator

    Moliboy5000/.claude/skills/skill-creator Skill

    no license0 repo

    Claude dotfile containing skills, agents etc I use.

  • Skill creator

    morzecrew/agent-skills/.agents/skills/skill-creator Skill

    0 repo

    Collection of agent skills

  • Skill creator

    mustafayilmazart/kesif-claude-skills/skills/skill-creator Skill

    0 repo

    14 production-ready Claude Code skill — TTS, video, NotebookLM, YouTube, token optimization, prompt engineering

  • Rag eval guardrails

    NeuralMedic-DE/claude-skills/rag-eval-guardrails Skill

    0 repo

    Verified-gate Claude Code skills that run their checks and prove the result — by NeuralMedic. Web, accessibility, healthcare & compliance, data, IaC, Odoo.

  • Pm trace eval

    Nyota-tree/pm-trace-eval/skills/pm-trace-eval Skill

    25 days old0 repo

    PM-facing AI conversation trace eval skill: north-star metrics, 1-5 scores, red/yellow/green, per-turn quote review

  • Ve skill benchmark adapter

    onyx679/automotive-ve-ai-skills-kit/skills/ve-skill-benchmark-adapter Skill

    0 repo

    AI Skill templates and workflow scoring tools for automotive value engineering and VAVE productivity scenarios

  • Eval refine

    pagoda111king/claude-code-pack/template/.claude/skills/eval-refine Skill

    no license0 repo

    Claude Code config pack · 11 agents · 10 skills · 10 commands · 4 hooks · CLAUDE.md template · Solo Founder OS

  • Idea eval

    pagoda111king/claude-code-pack/template/.claude/skills/idea-eval Skill

    no license0 repo

    Claude Code config pack · 11 agents · 10 skills · 10 commands · 4 hooks · CLAUDE.md template · Solo Founder OS

  • Add skill

    Paldom/landing-page-builder-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for building high-converting, modern, performant landing pages: design tokens, conversion copywriting, marketing UX, layout patterns, Next.js/React + shadcn implementation, scroll motion, and performance/SEO/accessibility.

  • Add skill

    Paldom/llm-council-skills/.claude/skills/add-skill Skill

    0 repo

    Agent Skills for building and operating LLM councils - multi-model deliberation with anonymized peer review, robust aggregation, calibrated escalation, and defenses against correlated failure.

  • Add skill

    Paldom/mcp-server-skills/.claude/skills/add-skill Skill

    17 days old0 repo

    Agent Skills for designing and implementing Model Context Protocol (MCP) servers - server architecture, tools, resources, prompts, transports, authorization, testing, and deployment.

  • Add skill

    Paldom/noslop/.claude/skills/add-skill Skill

    17 days old0 repo

    Agent Skills that strip AI writing tells from copy - em dashes, 'it's not X, it's Y' constructions, hedging, and bloat - so text reads like a human wrote it, with evals to verify.

  • Add skill

    Paldom/promptimize/.claude/skills/add-skill Skill

    18 days old0 repo

    Agent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.

  • Evaluate prompts

    Paldom/promptimize/skills/evaluate-prompts Skill

    18 days old0 repo

    Agent Skills for prompt optimization - tune prompts for the latest frontier models like Fable 5 and GPT 5.6, write effective goals and success criteria, and engineer agentic loops that converge.

  • Add skill

    Paldom/screenshooter/.claude/skills/add-skill Skill

    16 days old0 repo

    Skills for creating polished screen-recording walkthrough and tutorial videos of web apps, with smooth scripted mouse movements, zoom effects, and explanation cards.

  • Add skill

    Paldom/stealth-browser-skills/.claude/skills/add-skill Skill

    17 days old0 repo

    Agent Skills for stealthy, human-like browser automation with Playwright and Camoufox — real-browser agentic browsing that evades bot detection.

  • Add skill

    Paldom/terminaltor/.claude/skills/add-skill Skill

    15 days old0 repo

    Agent Skills for recording, sanitizing, and rendering terminal session demos - capture real command walkthroughs with agents, redact sensitive output, and render polished casts for READMEs and docs.

  • Ecc eval harness

    Pdbjork/hermosskills-site/skills/ecc-eval-harness Skill

    29 days old0 repo

    Hermesskills — a curated marketplace & repository of AI-agent skills. OpenClaw-compatible catalog at /catalog/v1/skills.json. 30% of commissions fund public-good AI.

  • Evaluation lm evaluation harness

    percymcn/agent-cookbook/skills/mlops/evaluation-lm-evaluation-harness Skill

    0 repo

    Production-ready AI agent skills, playbooks, and workflows from the pharma6 automation lab

  • Creating skills

    pgoell/pgoell-claude-tools/plugins/agent-system-management/skills/creating-skills Skill

    0 repo

    Collection of my personal claude skills

  • Benchmark models

    planifest/planifest-framework/planifest-framework/external-skills/benchmark-models Skill

    0 repo

    A specification framework for agentic development. Agents build from complete specs - not guesses.

  • Benchmark

    planifest/planifest-framework/planifest-framework/external-skills/benchmark Skill

    0 repo

    A specification framework for agentic development. Agents build from complete specs - not guesses.

  • Ai pm

    PlevanTem/luban-skill/.claude/skills/ai-pm Skill

    0 repo

    不是泛泛而谈,蒸馏一个专家,带走ta的方法论和专业判断。Distill expert methodology into agent kits — not vibes persona prompts

  • Skill creator

    ploteddie-bit/skills/skill-creator Skill

    no license0 repo

    Collection de compétences agentiques modulables pour Kimi, Aegis et assistants IA

  • Skill eval pipeline

    pngdeity/apm-user-repository/packages/skill-eval-pipeline/.apm/skills/skill-eval-pipeline Skill

    0 repo

    Personal APM marketplace — skills, prompts, agents, and instructions for AI coding agents

  • Llm wiki eval

    po4yka/llm-wiki-skills/skills/llm-wiki-eval Skill

    0 repo

    Portable Agent Skills for building, operating, evaluating, and governing LLM-Wiki knowledge systems.

  • Llm wiki eval tooling

    po4yka/llm-wiki-skills/skills/llm-wiki-eval-tooling Skill

    0 repo

    Portable Agent Skills for building, operating, evaluating, and governing LLM-Wiki knowledge systems.

  • Ai forge eval

    robcsaszar/ai-forge/skills/ai-forge-eval Skill

    26 days old0 repo

    Claude Code skills for creating, judging, evaluating, and updating skills and agents

  • Eval set

    robdasi/skills/eval-set Skill

    0 repo

    Free, working Claude skills I use to run an AI automation studio. Drop-in SKILL.md files. By Robin Laires / Laires Labs.

  • Eval log

    rulebased-io/claude-plugin/packages/harness/skills/eval-log Skill

    0 repo

    Claude Code plugins by rulebased.io — harness audit/init/recommend + second-brain capture/connect/review/organize

  • Eval log

    rulebased-io/harness/rulebased-harness/skills/eval-log Skill

    0 repo

    Harness engineering tools for AI agents — audit, init, recommend your project's agent-readiness (AGENTS.md, context engineering, eval)

  • Eval grader

    SamyakJhaveri/loam/seed/_research/skills/eval-grader Skill

    0 repo

    Copier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.

  • Eval run

    SamyakJhaveri/loam/seed/_research/skills/eval-run Skill

    0 repo

    Copier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.

  • Overnight eval

    SamyakJhaveri/loam/seed/_research/skills/overnight-eval Skill

    0 repo

    Copier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.

  • Post eval

    SamyakJhaveri/loam/seed/_research/skills/post-eval Skill

    0 repo

    Copier template that bootstraps AI-agent-optimized project setups: layered context routing, curated skills, and an enforced validation gate — synced across projects.

  • Amortized intelligence

    Sapien3/amortized-intelligence Skill

    0

    A Claude Code skill: spend frontier-model tokens at design time, compile them into artifacts — harnesses, rubrics, evals, skills — that cheap models run at near-zero marginal cost.

  • Ai rmf governor

    satishTheLegend/ai-rmf-governor Skill

    0

    Claude Code skill: run NIST AI RMF (Govern/Map/Measure/Manage) on your LLM/ML feature — named failure modes, metric thresholds, an offline eval gate, and signed-off residual risk.

  • Eval harness architect

    satishTheLegend/eval-harness-architect Skill

    0

    Stand up a rigorous, regression-proof evaluation harness for any LLM/agent system from zero — datasets, scorers, CI gates, and drift monitoring.

  • Instinct learn eval

    search-atlas-group/amm-founding-circle/skills/instinct-learn-eval Skill

    0 repo

    The AMM founding-circle home base: 36 Claude skills (AEO/SEO + agentic engineering + security), the agentic ladder, playbooks, and automations. No paid APIs required.

  • Agent eval

    Sheshiyer/skill-clusters/skills/agent-eval Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Agentic engineering

    Sheshiyer/skill-clusters/skills/agentic-engineering Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Benchmark

    Sheshiyer/skill-clusters/skills/benchmark Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Eval harness

    Sheshiyer/skill-clusters/skills/eval-harness Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Evals

    Sheshiyer/skill-clusters/skills/evals Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Healthcare eval harness

    Sheshiyer/skill-clusters/skills/healthcare-eval-harness Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Hugging face community evals

    Sheshiyer/skill-clusters/skills/hugging-face-community-evals Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Optimize

    Sheshiyer/skill-clusters/skills/optimize Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Quality eval core

    Sheshiyer/skill-clusters/skills/quality-eval-core Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Quality eval orchestrator

    Sheshiyer/skill-clusters/skills/quality-eval-orchestrator Skill

    0 repo

    Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

  • Ai eval regression tester

    sisodiabhumca/agent-skills/skills/ai-eval-regression-tester Skill

    0 repo

    Production-Ready Agent Skills : product analytics, growth experiments, CRM, research synthesis, postmortems, data contracts, SaaS spend, compliance, architecture maps, and LLM eval and many more.

  • Llm evaluation

    SkillMedev/ai-engineer-toolkit/skills/llm-evaluation Skill

    0 repo

    Ship LLM features that work — prompts, evals, agents, and MCP servers.

  • Eval authoring

    soltraveler-sri/Steerable-Protocol/skills/eval-authoring Skill

    29 days old0 repo

    An open standard for making apps steerable, shipped with a reference runtime and a skill that lets a coding agent implement it in your codebase.

  • Skill creator

    songsunny00/MySkills/BaseSkills/.agents/skills/skill-creator Skill

    no license0 repo

    根据不同角色,收集或编写日常常用skill

  • Skill creator

    sorinLupaIonut/digital-fte/.claude/skills/skill-creator Skill

    15 days old0 repo

    A customer-support Digital FTE (AI Worker): OpenAI Agents SDK on a sandboxed runtime, capabilities as portable Skills, a scoped customer-data MCP server instead of raw SQL, pgvector semantic search on Neon Postgres, and a transactional audit row behind every action.

  • Cross eval

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Business & Ops/cross-eval Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.

  • Eval

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Data & AI/eval Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.