agentsclimarketplace

Evals & benchmarks

946 rows, by stacks then stars

  • Skill creator

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Development/skill-creator Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.

  • Hugging face community evals

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Product/hugging-face-community-evals Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.

  • Self eval

    srg-sphynx/MDForge/Sources/MDForge/Resources/SkillLibrary/Research/self-eval Skill

    0 repo

    Native macOS skill-catalog studio for Claude Code — browse 1,492 Markdown skills, customize with variables or AI (hosted or local), export to .claude/skills. SwiftUI + Liquid Glass.

  • Skill factory

    srinitude/skills/skills/skill-factory Skill

    16 days old0 repo

    All of my portable Agent Skills with local validation and MCP access

  • Image benchmark report

    StaryMoon/image-benchmark-report-skill/skills/image-benchmark-report Skill

    23 days old0 repo

    Agent Skill for PSNR/SSIM image benchmarks with per-image scores and worst-case reports.

  • Stayingapi comp set benchmark

    stayingapi/travel-skills/skills/stayingapi-comp-set-benchmark Skill

    0 repo

    Portable AI-agent skills (SKILL.md) for accommodation search, availability, pricing, cross-OTA comparison & reviews via StayingAPI. Works with Claude, MCP and more. MIT-0.

  • Skill builder

    Stoica-Mihai/claude-skills/plugins/skill-builder/skills/skill-builder Skill

    0 repo

    Curated Claude Code plugin marketplace — OpenSpec extensions and autonomous development workflows

  • Eval engineer

    sumitake/agent-collab/plugins/agent-collab/skills/eval-engineer Skill

    24 days old0 repo

    Unified dynamic-host agent collaboration policy and signed-runtime client

  • Hallucination investigator

    sumitake/agent-collab/plugins/agent-collab/skills/hallucination-investigator Skill

    24 days old0 repo

    Unified dynamic-host agent collaboration policy and signed-runtime client

  • Ai evaluation

    SWEStash/swe-workflow-skills/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/ai/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/data-scientist/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Ai evaluation

    SWEStash/swe-workflow-skills/plugins/ml/skills/ai-evaluation Skill

    0 repo

    A comprehensive SWE workflow, encoded. Might be useful to you too.

  • Skill evaluator

    takehiro177/skill-evaluator/skills/skill-evaluator Skill

    0 repo

    Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

  • Skill benchmark

    TheophilusChinomona/idev/skills/skill-benchmark Skill

    0 repo

    Claude Code plugin: token-optimized dev workflow — cached project context, pattern scanners, build & wiring verification, session persistence, self-review. 23 skills, 3 agents, 4 commands.

  • Skill creator

    ThiagoPanini/aidriven/.github/skills/skill-creator Skill

    0 repo

    🔩 A Python CLI tool for getting and installing AI development resources into software projects.

  • Skill creator

    ThiagoPanini/aidriven/.claude/skills/skill-creator Skill

    0 repo

    🔩 A Python CLI tool for getting and installing AI development resources into software projects.

  • Traigent eval audit

    Traigent/traigent-skills/skills/traigent-eval-audit Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Traigent eval build

    Traigent/traigent-skills/skills/traigent-eval-build Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Traigent eval choose metric

    Traigent/traigent-skills/skills/traigent-eval-choose-metric Skill

    no license0 repo

    Agent skills for deploy and use Traigent

  • Vezir yetistirme

    trugurpala/divan/plugins/sadrazam/skills/vezir-yetistirme Skill

    20 days old0 repo

    Divan — vibe coder vezirler kurulu. Claude Code / Cursor / Codex skill paketleri: Sadrazam (uctan uca orkestrator), planlama, TDD, debugging, UI/UX, React.

  • Library eval

    turntuptechnologies-ai/skills/skills/library-eval Skill

    0 repo

    Turnt Up Technologies株式会社の Claude Code プラグインマーケットプレイス

  • Anti hallucination

    udsy19/.claude/skills/anti-hallucination Skill

    no license0 repo

    A production-grade .claude folder for Claude Code: 40 lifecycle skills, agents, commands, and hooks — with reflexive skill routing and always-on engineering disciplines (verify, no-bloat, continuous-git, memory, ship-fast).

  • Cis benchmark audit

    vahagn-madatyan/netsec-skills-suite/skills/cis-benchmark-audit Skill

    0 repo

    The Complete Network & Network Security Skills Suite

  • Authoring skills with evals

    vemodalen-x/VEMO_SKILLS/skills/governance/authoring-skills-with-evals Skill

    0 repo

    Public reusable skill hub for agent workflows

  • Rendering html eval reports

    vemodalen-x/VEMO_SKILLS/skills/orchestration/rendering-html-eval-reports Skill

    0 repo

    Public reusable skill hub for agent workflows

  • Negative skill eval writer

    vibesec-advisory/vibesec-advisory-skill-library/SKILLS/ai-workflow-eval-and-lifecycle-review/skills/negative-skill-eval-writer Skill

    0 repo

    Free AI workflow skill libraries for GTM teams, with implementation patterns, guardrails, and evals.

  • Eval case converter

    vibesec-advisory/vibesec-advisory-skill-library/SKILLS/workflow-failure-label-review/skills/eval-case-converter Skill

    0 repo

    Free AI workflow skill libraries for GTM teams, with implementation patterns, guardrails, and evals.

  • Behavior spec canvas

    VJDiPaola/skill-forge/behavior-spec-canvas Skill

    11 days old0 repo

    A skill library for AI coding agents with a CI quality gate: 38 skills, four-target sync, and an 18-check eval harness that fails the build on malformed skills.

  • Eval framework

    VJDiPaola/skill-forge/eval-framework Skill

    11 days old0 repo

    A skill library for AI coding agents with a CI quality gate: 38 skills, four-target sync, and an 18-check eval harness that fails the build on malformed skills.

  • Skill creator

    VRIL-LABS/skill-jam/skills/ai-ml/claude-plugins-official-main/claude-plugins-official-main/plugins/skill-creator/skills/skill-creator Skill

    no license0 repo

    Welcome to the skill-jam ☄️🏀

  • Llm observability and evals

    wanghong5233/agent-engineering-kit/cursor/.cursor/skills/llm-observability-and-evals Skill

    0 repo

    A reusable, production-grade .cursor/ engineering package for Cursor / Claude Code / Agent IDEs. Rules, skills, commands, and deterministic safety hooks extracted from a real Agent project.

  • Skill creator

    Waterinsoluble-orderplatyctenea746/book2skills/skills/skill-creator Skill

    0 repo

    Transform classic books into AI agent skills for practical, source-based decision making

  • Skill creator

    wentallout/write-like-a-human/.agents/skills/skill-creator Skill

    no license0 repo

    Messy, opinionated, and aggressively human.

  • Ai eval harness

    willianbs/skills/ai-eval-harness Skill

    19 days old0 repo

    AI Engineering Operating System

  • Audit the oracle coverage

    XyraSinclair/ideonomy/skills/audit-the-oracle-coverage Skill

    0 repo

    Computational ideonomy: Gunkel's science of ideas as inference-time machinery — a 37-primitive organon, an MDL-ratcheted respiratory engine, and 14 gated agent skills (Claude Code plugin)

  • Skill creator

    yanickfischer/zhaw-bsc-latex-claude-skelett/.claude/skills/skill-creator Skill

    no license0 repo

    ZHAW BSc-Thesis-Vorlage (Wirtschaftsinformatik) — LaTeX-Skelett mit XeLaTeX, APA-7 via biblatex, Zotero-Pipeline, Draft/Submission-Toggle. Erweitert mit Claude-Code-Skills für ZHAW-Formalia (KI-Deklaration, wiss. Schreiben, Layout-Check, Mermaid-Diagramme).

  • Clinical note eval

    yashpatil582/agent-skills/clinical-note-eval Skill

    0 repo

    Anthropic Agent Skills packaging shipped eval methods: gaming-resistant clinical-note grounding eval + token-F1 scorer. Stdlib-only, runs offline.

  • Token eval harness

    yashpatil582/agent-skills/token-eval-harness Skill

    0 repo

    Anthropic Agent Skills packaging shipped eval methods: gaming-resistant clinical-note grounding eval + token-F1 scorer. Stdlib-only, runs offline.

  • Drift eval

    Yco-0314/strata/skills/l5-meta/drift-eval Skill

    0 repo

    Strata — a self-measuring skills framework for Claude Code & Codex: best-of-breed engineering skills stacked by altitude, where every skill has to prove it moves a measured number.

  • Skill eval

    Yco-0314/strata/skills/l5-meta/skill-eval Skill

    0 repo

    Strata — a self-measuring skills framework for Claude Code & Codex: best-of-breed engineering skills stacked by altitude, where every skill has to prove it moves a measured number.

  • Eval audit and sweep

    YosukeIida/personal-agent-skills/eval-audit-and-sweep Skill

    20 days old0 repo

    Personal agent skills (Claude Code / Codex) — SKILL.md collection

  • Prompt rail

    yuanGao0816/prompt-rail Skill

    16 days old0

    Measured prompt iteration with train/test dual-gate scoring and anti-overfit rails. Agent Skill (SKILL.md).

  • Wechat benchmark research

    yuanjiemi/yuanjiemi-ai-skills/skills/wechat-benchmark-research Skill

    18 days old0 repo

    YUAN元解密 AI Skills:面向真实内容运营场景的可安装智能体技能

  • Skill creator

    zhijunio/skills/skill-creator Skill

    0 repo

    个人 Agent Skills:健康、阅读、写作、日记、Keep 等脚本与说明

  • Buyer eval skill

    Zorineinsupportable217/buyer-eval-skill Skill

    0

    Evaluate B2B vendors by talking to AI agents and verifying claims against independent sources to score what is true and what is not