agentsclimarketplace

Auditing skills

Skill AdrianParedez/capability-fabric/skills/auditing-skills

A model-agnostic Agent Skills library for explicit routing, bounded context, and verifiable agent behaviour.

Install
npx -y skills add AdrianParedez/capability-fabric --skill auditing-skills

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Scores and improves an existing Agent Skill against a spec + quality rubric, then proposes or applies concrete refactors (split monoliths, sharpen descriptions, fix portability). Use to review, grade, lint, or improve a skill, or to maintain a skill library at scale. Justifies each finding with verifying-reasoning and routes fixes back to authoring-skills.

The file declares its own license as Apache-2.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

4.3 KB, as published. Nobody here has run it

Auditing skills

Authoring creates skills; auditing keeps them good as the library grows. This skill grades a skill against an objective rubric, finds the highest-leverage problems, and turns them into concrete fixes, so quality scales past what any single author tracks.

Use this when

  • Reviewing / grading / linting / improving a skill (yours or third-party).
  • Maintaining a library of many skills (periodic audit, pre-merge gate).
  • After authoring-skills produces a skill (the independent check).
  • Skip for a quick personal note-to-self skill no one else uses.

The audit

- [ ] 1 VALIDATE spec compliance (name/description/fields), pass/fail, hard gate
- [ ] 2 SCORE    rate each rubric dimension 0-2 (see references); sum to a grade
- [ ] 3 DIAGNOSE find the top 1-3 issues by leverage (what most hurts use/cost/portability)
- [ ] 4 JUSTIFY  for each finding, give evidence (a line, a missing case), verifying-reasoning
- [ ] 5 FIX      propose concrete edits; for big issues, route to authoring-skills (e.g. split)
- [ ] 6 RECHECK  re-score after fixes; confirm evals still/now pass
- [ ] 7 REPORT   grade + findings + diffs + before/after score

The rubric (0-2 each; see references for anchors)

#DimensionAsks
R1Spec validityfrontmatter rules met? (hard gate, fail = blocked)
R2Description qualitywhat+when, third person, distinct from neighbors, trigger words
R3Single responsibilityone capability? any "and also" smell?
R4Concisenessevery token justified? explains things the model knows?
R5Progressive disclosurebody <500 lines? detail split, one level deep?
R6Activation safety"use when/not when"? overlaps disambiguated?
R7Portabilitycapabilities not tool names? fwd slashes? runtime-adaptation?
R8Evidence/examplesconcrete examples + anti-example? consistent terms?
R9Testability≥3 evals present? benchmarks measurable?
R10Safety/scriptssolve-don't-punt, no magic constants, deps declared, one-way gated

Grade = Σ/20. Ship ≥16 and R1 pass and no dimension at 0.

Diagnose by leverage (don't list everything)

Rank issues by impact, not count. A vague description (R2) that prevents activation wastes the whole skill, fix that before a cosmetic wording nit. The audit's value is prioritized improvement, not an exhaustive lint dump.

Common fixes the audit prescribes

FindingFix
Monolith (R3 low)split by capability; route via composing-skills
Vague description (R2 low)rewrite what+when, add trigger keywords, disambiguate
Bloated body (R4/R5 low)cut model-known content; move depth to references/
Tool-name coupling (R7 low)replace with capability phrasing + runtime-adaptation note
No evals (R9 low)generate ≥3 from observed gaps (authoring-skills step 2)
Overlap misfire (R6 low)add "for X use Y" cross-links to both skills

Apply vs propose

  • Propose (default for third-party / high-risk): output the findings + diffs for review.
  • Apply (when authorized & low-risk): make the edits, then RECHECK. Treat edits to a shared skill as a change requiring the guarding-code-quality honesty in the report.

Runtime adaptation

  • Minimum: filesystem read. Use skills-ref validate for R1 if available; else apply the name/description rules manually.
  • Auditing is reasoning + file reads, portable across all runtimes.

Files

  • references.md, rubric anchors (what 0/1/2 looks like per dimension).
  • examples.md, a real audit with score, findings, and fixes.
  • templates/audit-report.md, the report format.
  • checklists/audit.md, the audit pass as a checklist.
  • benchmarks/, does the audit catch seeded defects and improve scores?

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.