agentsclimarketplace

Skillopt

Skill skillberry-ai/cap-evolve/skills/algorithms/skillopt

Optimize any AI agent’s skills, tools/MCP, and prompts against your own evals.

Install
npx -y skills add skillberry-ai/cap-evolve --skill skillopt

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Runs the SkillOpt single-lineage optimization loop over epochs x mini-batches with a textual learning rate (an integer edit budget that decays on a constant|linear|cosine schedule), a within-epoch rejected-edit + failure-pattern buffer injected into the optimizer prompt, and a gated epoch-boundary slow/meta update that fixes longitudinal regressions. Parent is always the current best; acceptance is the val significance gate; the test split stays sealed. Use when a run benefits from disciplined annealing and per-epoch consolidation rather than hill-climb's one-shot whole-trainset proposals or gepa's Pareto frontier.

SKILL.md

5.3 KB, as published. Nobody here has run it

skillopt — annealed single-lineage climb (epochs × mini-batches)

SkillOpt (arXiv:2605.23904) borrows the deep-learning training loop. It is a strict single-lineage climber (parent is always the current best, like hill-climb) but adds three things from the DL analogy:

  • a textual learning rate — an integer edit budget L per step, large early (explore broadly) and shrinking later (consolidate), on a constant | linear | cosine schedule;
  • a per-epoch rejected-edit + failure-pattern buffer injected into the next step's prompt (don't re-propose dead ends; these failures remain unsolved);
  • an epoch-boundary slow / meta update: re-evaluate the epoch-start skill vs the current best on a small train sample, categorize improved / regressed / persistent / stable, then run ONE extra gated step to fix regressions without breaking stable successes.

Every step calls the shared run_step (materialize → optimize → eval-VAL → significance gate → accept/reject → snapshot/best, with RejectedMemory/History); this skill only owns the schedule, the buffer, and the slow update. Gated on val; test stays sealed (that's finalize).

The DL analogy

deep learningSkillOpt
epoch over the datasetepoch over the train ids (shuffled, seeded by epoch)
mini-batch / gradient accumulation--batch-size train tasks × --accumulation per step
learning-rate scheduleedit-budget L schedule (--lr-schedule, decays --edit-budget--min-edit-budget)
gradient stepone run_step (a bounded edit, gated on val)
momentum / replaythe per-epoch rejected-edit + failure-pattern buffer
LR warm-restart / fine-tunethe epoch-boundary gated slow/meta update

L is communicated to the optimizer in natural language only ("make at most L bounded edits") — the LLM is not mechanically clipped — so the loop logs requested-vs-applied edits to surface an optimizer that ignores its budget.

When to use vs hill-climb / gepa

algorithmparentproposal unitbest when
hill-climbcurrent bestwhole train set, one-shot each iterbroad gaps; you want the simplest loop
skilloptcurrent best (single lineage)mini-batch under a shrinking edit budget, epochs + gated slow updateyou want disciplined annealing + per-epoch consolidation; a moderately sized train set
gepasampled from a per-instance Pareto frontierminibatch with a cheap local gate before full valmany specialists the mean hides; merge across lineages

Inputs / outputs (manifest tokens)

  • needs: scores + traces (per-task val results to reflect on) and candidate (the parent to extend).
  • provides: candidate (the accepted best).

Standalone use

python scripts/run.py --run-dir .capevolve/run_X --project .capevolve/project \
  --optimizer 'python .../run-optimizer/scripts/run.py --name mock --workdir {workdir} --prompt {prompt}' \
  --epochs 4 --batch-size 8 --edit-budget 4 --lr-schedule cosine --min-edit-budget 2 \
  --n-trials 4

Requires baseline.json first (like its sibling algorithms). --resume continues from the run's current best. --no-slow-update disables the epoch-boundary meta step; --no-regression adds a SWE-bench-style dual gate.

Pitfall: small val sets

The default gate is significant/paired (NOT naive strict-greater) so a tiny val set does not reject every edit on noise. With a small val set, raise --n-trials (real per-trial variance) or use a graded reward so the paired significance test has signal. Only the slow update is a "meta" step — it is still gated on val, never force-accepted, and its train sample is small + counted in budget (toggle with --no-slow-update).

Agent-mode loop

When orchestration_mode: agent, drive skillopt yourself: single-lineage climb with a textual learning rate over epochs/minibatches — each step propose an edit sized to the current LR, evaluate on val (minibatch then full as skillopt prescribes) via cap-evolve, gate Δ>k·SE, accept→snapshot / reject→revert & decay. Log steps to the run dir; between steps verify rollouts+results landed for the dashboard. Re-read stop_condition; stop on it/stall/budget. Seal once with cap-evolve finalize, then report.

References

  • references/concepts.md — the SkillOpt loop in detail, the textual-LR schedule, the buffer/slow-update mechanics, and citations (arXiv:2605.23904).

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.