Skill benchmark
Claude Code plugin: token-optimized dev workflow — cached project context, pattern scanners, build & wiring verification, session persistence, self-review. 23 skills, 3 agents, 4 commands.
npx -y skills add TheophilusChinomona/idev --skill skill-benchmarkAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Evaluates and benchmarks Agent Skill quality via static analysis and evaluation methodology. Use when assessing or reviewing skill quality, benchmarking skills, measuring skill activation accuracy, optimizing skill descriptions, or running /idev:benchmark-skills.
SKILL.md
3.5 KB, as published. Nobody here has run it
Skill Benchmark — Evaluate Agent Skill Quality
Evaluate skills through static analysis (automated, run the script) and evaluation-driven methodology (manual, follow the process below). Adapted from Anthropic's skill evaluation guidance.
Static analysis (automated)
Run the checker against any plugin:
python3 "${CLAUDE_PLUGIN_ROOT}/skills/skill-benchmark/benchmark_skills.py" # this plugin
python3 "${CLAUDE_PLUGIN_ROOT}/skills/skill-benchmark/benchmark_skills.py" --plugin-dir <dir>
It scores every skills/*/SKILL.md against these checks:
| Check | Pass criteria |
|---|---|
| desc-present | Description non-empty, ≤ 1024 chars |
| desc-triggers | Contains "Use when/before/after/for" activation triggers |
| desc-3rd-person | No "I can" / "You can" phrasing |
| name-kebab | Matches ^[a-z0-9]+(-[a-z0-9]+)*$ |
| name-length | ≤ 64 chars |
| name-reserved | No "anthropic"/"claude" in the name |
| name-matches-dir | Frontmatter name equals directory name |
| body-length | SKILL.md ≤ 500 lines |
| has-examples | Contains at least one fenced code block |
| refs-resolve | Every ${CLAUDE_PLUGIN_ROOT}/... path in the file exists |
--strict --min-score N makes it exit non-zero below a threshold (CI use).
Skill categories
Classify each skill before evaluating — the test differs:
- Capability uplift: enhances core abilities (analysis, verification, code patterns). Stable across model versions. Test by comparing base-model output with and without the skill.
- Encoded preference: encodes project/team workflows and conventions (commit formats, cache layouts). Test by verifying fidelity to the encoded workflow, not output quality.
Evaluation methodology (manual)
For activation accuracy and effectiveness, static checks aren't enough:
- Write evals per skill: 5-10 representative prompts that SHOULD trigger
it, 3-5 out-of-scope prompts that should NOT, and expected-behavior
criteria for each. Use
evaluation-checklist.mdin this skill's directory as the template. - A/B test: run agent A with the skill loaded and agent B without; have a third agent judge outputs blind. Track pass rate, token usage, time.
- Multi-model targets: Haiku 70%+, Sonnet 85%+, Opus 95%+ pass rate. If only Haiku fails, the instructions rely on implicit reasoning — make them more explicit.
Description optimization
The description alone determines activation. Tune for ~90%+ true positives and <5% false positives:
- Too many false positives → description too broad: add domain-specific terms.
- Too many false negatives → too narrow: add synonyms and adjacent phrasings.
Iteration
Test → measure → analyze → refine → verify. Stop when eval pass rates meet the model-tier targets, activation is accurate, and token cost is justified by the benefit.
Reporting rules (anti-fabrication)
- Run the script (or the evals) BEFORE making any quality claim; report only what was measured.
- Never invent percentages, pass rates, or activation numbers — if unmeasured, say "not yet measured".
- No superlatives ("excellent", "comprehensive", "optimal") in assessments; state which checks passed and failed.
- Verify files exist (Read/Glob) before claiming they do.