Add skill
From context to installable agent skills - research packs in, validated skills.sh-ready skills out. Scaffolding, eval-first authoring, validation, and deployment included.
npx -y skills add Paldom/skillskit --skill add-skillAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Authors or improves an Agent Skill in the current repo via the eval-first workflow - scope one purpose, write trigger evals, draft SKILL.md, validate, benchmark. Use when the user asks to create, add, write, refine, fix, or benchmark a skill ("add a skill for X", "improve triggering of Y"). Not for scaffolding a new repo (create-skill-repo) or distilling research packs (skill-from-research).
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
5.8 KB, as published. Nobody here has run it
add-skill
Builds one production-quality skill in skills/<name>/, eval-first, in whatever
repository you are currently in. Read references/skill-authoring.md and
references/evals.md once per session before starting — they are the rulebook
this workflow enforces (repos scaffolded by create-skill-repo also carry them
as docs/). If invoked as /add-skill <args>, treat $ARGUMENTS as the skill
name or idea to scope in step 1.
When NOT to use
- Creating a new skills repository →
create-skill-repo. - Turning a research pack into one or more skills →
skill-from-research(it drives this workflow per skill after distilling the pack). - Infrastructure work (hooks, CI, validators) → plain editing, PR-level scrutiny.
- A "skill" that merely restates what the model already does well → don't build it; a skill must fix a specific observed failure.
Workflow
- Scope. State in one sentence what the skill does and the concrete failure it fixes. If the sentence needs "and", split into multiple skills, one at a time. Check the repo's existing skill catalog for overlap — near-neighbor descriptions steal each other's triggers; plan disjoint wording.
- Gather. Read the repo's
.local/recursively if present (research, examples, constraints — never cite.local/paths from a skill). Research beyond it: web-search for current facts, official docs, and prior art; verify anything load-bearing against primary sources. Verified facts the skill depends on go intoskills/<name>/references/as cleaned, committed files. - Evals first. Create
skills/<name>/evals/evals.jsonperreferences/evals.md: ≥8should_trigger(vary formality, typos, terseness), ≥8should_not_trigger(near-misses sharing keywords), 3–5qualitycases with plain-languageexpected_behaviorassertions. If you can't write these, the scope is unclear — back to step 1. - Draft SKILL.md. Frontmatter:
name== folder name;descriptionon a single line, third person,[what] + [use when …] + [not for …], 150–400 chars, trigger keywords in the first ~120. Body perreferences/skill-authoring.md: purpose → when/when-not → numbered workflow → output spec → gotchas → pointers. Under 500 lines; deterministic steps asscripts/(non-zero exit on failure); long material asreferences/with a TOC. Reference bundled files via${CLAUDE_SKILL_DIR}so installed copies work. - Validate. If the repo has a validator (
make check/scripts/validate_skills.py), it must exit 0 — fix every error, triage every warning. If it has none, self-check against the frontmatter rules in the authoring reference (single-line description, name==folder, length bounds). - Trigger self-test. For ≥3 should-trigger and ≥2 should-not-trigger prompts, reason explicitly: would the description alone route this prompt here against every other installed skill? Fix the description, not the evals.
- Measure. Scan the available-skills list for Anthropic's official
skill-creator(typicallyskill-creator:skill-creator; the plugin-qualified name varies by marketplace — match on name/description, don't hardcode). If present, invoke it as a subskill via the Skill tool, scoped to measure only: it runs each quality case as with-skill vs baseline subagents, grades assertions, and benchmarks pass rate / time / tokens. Hand it cases fromevals/evals.jsontranslated per "Automated harness" inreferences/evals.md, and follow that section's guardrails (workspace-only writes, cases-run check, description acceptance gate). If it's absent, run the manual protocol instead and say so — it certifies triggering only, not quality uplift; never let a skipped Measure pass silently as "measured". - Register. Update the repo's README catalog,
CHANGELOG.md, andskills.sh.jsongrouping when those exist (schema shape:"groupings": [{"title": ..., "description": ..., "skills": [...]}]— notgroups/name; skills.sh silently ignores a non-conforming file). Re-check sibling descriptions for new overlap.
Output spec (Definition of Done)
skills/<name>/with SKILL.md + evals/evals.json (+ scripts/references as needed)- Repo validator green (or the self-check documented when no validator exists)
- Measured: official skill-creator benchmark when installed, else the manual
protocol from
references/evals.md— state which ran (the fallback certifies triggering only) - Catalog/CHANGELOG/groupings updated where present
- A 3-line summary: purpose, trigger phrase examples, known limitations
Gotchas
- Keep
descriptionon ONE physical line — multi-line values have failed to load in some runtimes; never work around a repo's validator. - Don't paste failed eval prompts verbatim into the description (overfitting); generalize the missing verb/noun instead.
- Over-triggering fix = add targeted "Not for …" noun phrases from the actual false-positive overlap; don't delete positive triggers.
- Model upgrades shift routing — note that evals should be re-run after the next model release.
Files
references/skill-authoring.md— anatomy, frontmatter rules, description formula, body structure, gotchas, anti-patterns.references/evals.md— evals.json format and the trigger/quality protocol.