Optimize skill
Skill Bennyoooo/skillmaxxing/skill-maxing-plugin/skills/optimize-skill
Improve an existing skill through an evaluation-gated loop — run it on its eval tasks, analyze failures, propose bounded edits, validate a candidate, and (with your approval) promote a strictly-better version. Use when a skill underperforms or the user asks to optimize/tune/improve a skill that has an eval set.From its SKILL.md
npx -y skills add Bennyoooo/skillmaxxing --skill optimize-skillAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 21 stars21 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 5 commands, including `scripts/optimize.sh score --eval eval.yaml --rollouts rollouts.json --skill <name> --json` and 4 more.
SKILL.md
3.6 KB, 753 tokens by cl100k_base, as published. Nobody here has run it
optimize-skill
Make a skill measurably better without uncontrolled drift. This is an agent-in-the-loop loop, not a hands-off run: the CLI owns the deterministic machinery (scoring, edit budget, rejected-edit buffer, the gate, atomic promote/revert); you own the reasoning (running the skill, judging prose outputs, proposing edits). Expect several turns per optimization.
Preconditions
- The skill has an eval manifest (
eval.yaml) with real tasks. If it has none, stop and offer to create one (create-skill) — optimization cannot run without an eval set. - Optimization edits a managed copy, never the installed symlink target.
The loop (repeat until the gate stops improving)
-
Rollout. For each eval task
input, run the current skill yourself and collect its output. Writerollouts.json:[{ "taskId": "...", "output": "..." }]. -
Score.
scripts/optimize.sh score --eval eval.yaml --rollouts rollouts.json --skill <name> --jsonDeterministic tasks are scored for you.
agent-judgetasks come back as pending — score those yourself against each task's rubric and fold them into the aggregate. Record the current score. -
Reflect. Read the failing trajectories. Diagnose why they failed (this is your job — the CLI never judges why). Propose a small set of structured edits to
SKILL.md. Writeedits.json: an array of{ op: append|insert_after|replace|delete, target?, content?, sourceType: "failure"|"success", supportCount? }. Prefer failure-driven edits. -
Apply (bounded).
scripts/optimize.sh apply --skill <name> --skill-dir <live-dir> --edits edits.json --step <n> --total <N>The CLI caps edits at the budget (annealed over steps), skips edits to the protected
SLOW_UPDATEregion, and writes a candidate copy. Note the candidate dir it prints. -
Validate. Re-run rollout + score against the candidate (steps 1–2 pointing at the candidate dir), including the held-out tasks. Then gate:
scripts/optimize.sh gate --current <currentScore> --candidate <candidateScore> --best <bestScore>A non-zero exit means reject — add those edits to your rejected set so you don't re-propose them, and try a different reflection. Also reject if any held-out task regressed, even if the aggregate improved.
-
Promote (human gate). Only on a strict improvement with no held-out regression, present the candidate and its score delta to the user. On their approval:
scripts/optimize.sh promote --skill <name> --live <live-dir> --candidate <candidate-dir> --score <candidateScore>The prior version is retained and the change is reversible.
Revert
scripts/optimize.sh revert --skill <name> --version <prior-version> --live <live-dir>
Honesty
- "Optimize automatically" means the loop, budget, buffer, and gates are automated — the intelligence (rollout, reflection, edits, agent-judge) is yours. A weak reasoning pass simply makes less progress; the gate guarantees no regression is ever promoted.
- Never promote without explicit user approval, even when the gate passes.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.