Optimize
A collection of personal AI coding assistant configurations, specialist agents, and automated workflows optimized for Python and ML open-source development.
npx -y skills add Borda/AI-Rig --skill optimizeAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 23 stars23 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Minimal codex-native optimization loop. Use for metric-driven improvements with guardrails and measurable gates.
SKILL.md
6.7 KB, ~1.5k tokens by cl100k_base, as published. Nobody here has run it
Optimize
Metric-driven optimization with explicit guards, rollback criteria, experiment log.
Input Schema
{
"goal": "required measurable improvement objective",
"mode": "single|campaign",
"metric_cmd": "required command that emits or validates the target metric",
"metric_direction": "higher|lower",
"guard_cmd": "required command that must continue to pass",
"max_iterations": "optional integer, default 1",
"min_delta": "optional practical significance threshold",
"scope_files": [
"paths the optimization may edit"
],
"done_when": "metric improves without guard regression"
}
Workflow
01: Create run directory
Run python PLUGIN_ROOT/shared/create_run.py --skill optimize once. Retain its single printed path as
<run-directory> and substitute that literal path into every later artifact path and helper argument. Never store or
reuse the path through a shell variable; shell variables do not persist across tool calls.
02: Validate metric and guard commands
Require:
- Repeatable
metric_cmdproducing comparable value or pass/fail. - Known
metric_direction. guard_cmdfails on unacceptable regressions.- Bounded
scope_files. - Explicit, bounded
max_iterationsforcampaign. - Protect files/scripts used by
metric_cmd/guard_cmdunless user explicitly scopes them and accepts measurement-integrity risk.
Dry-run both before edit:
Execute the configured metric_cmd and guard_cmd separately with the host-native command runner. Write complete
combined output to <run-directory>/metric-baseline.txt and <run-directory>/guard-baseline.txt; retain both exit
codes and stop before editing if either command cannot run.
03: Record baseline and hypothesis
Write <run-directory>/hypothesis.md:
- metric to improve
- expected mechanism
- files allowed to change
- guard risk
- rollback condition
For campaign, noisy metrics, GPU/ML performance, or correctness-sensitive code, apply ../../shared/specialist-orchestration.md. Write <run-directory>/specialist-optimization-plan.md with narrow context packs for:
squeezer: profiling mechanism, bottleneck hypothesis, measurement plan.qa-specialist: guard coverage and regression risk.data-steward: data pipeline or reproducibility impact.scientist: metric validity, ablation design, statistical noise.challenger: overfitting to the metric or weakening guard checks.
No fan-out for one small measured change with stable metric/guard. Never let specialist change metric/guard scripts unless explicitly in scope_files and measurement-integrity risk recorded.
Structural context (optional): when scope_files resolves to a Python module/symbol, also probe codemap-py once
for callers, coupling, and test impact before the first iteration: python PLUGIN_ROOT/shared/codemap_adapter.py context --category develop --target <qname> --out <run-directory>/codemap-context.json. Per
../../shared/codemap-contract.md, absence/incompatibility is non-fatal — continue with the hypothesis above.
Persist the result once here, before step 04 applies any change; any triggered specialist consumes
<run-directory>/codemap-context.json, never a fresh query.
Initialize machine-readable iteration log:
Create an empty <run-directory>/experiments.jsonl with the filesystem tool before the first iteration.
04: Apply one minimal optimization change per iteration
One independent hypothesis per iteration. Do not optimize unmeasured paths. Before each, write <run-directory>/iteration-<n>-before.patch with scoped-file diff. If iteration fails and only its patch is present, revert with git apply -R against iteration diff; otherwise fail run when clean reversal cannot be proven. Never use git reset --hard.
05: Re-measure
Re-run the same retained metric_cmd and guard_cmd separately with the host-native command runner. Write complete
combined output to <run-directory>/metric-after.txt and <run-directory>/guard-after.txt; retain both exit codes.
06: Compare baseline and after results in <run-directory>/comparison.md
Required fields:
- baseline value
- after value
- delta
- guard status
- confidence
- noise caveats
Append one JSON object/iteration to <run-directory>/experiments.jsonl:
{
"iteration": 1,
"hypothesis": "one-line mechanism",
"metric_before": 0.0,
"metric_after": 0.0,
"delta": 0.0,
"guard": "pass|fail",
"decision": "kept|reverted|inconclusive|failed",
"rollback_evidence": "path or reason"
}
07: Decide keep/revert
- Keep only with intended metric movement and passing guards.
- With
min_delta, keep only if delta meets/exceeds practical-significance threshold. - Revert or fail on guard regression.
- For noisy measurement, repeat or mark inconclusive.
- In
campaign, stop at first kept result unless user asked continued exploration; otherwise continue only whilemax_iterationsremains and each rejected iteration has rollback evidence.
08: Run shared quality gates
Inspect python PLUGIN_ROOT/shared/run_gates.py --help. Tests runs configured test or guard command; give real commands or explicit reasons for other gates.
09: Write and validate the mandatory result artifact
Follow ../../shared/helper-cli-contract.md and authoritative help. Write OPTIMIZE_METADATA, validate as optimize, promote only validated candidate.
Fail-Fast Rules
- Missing metric or guard command => fail.
- Baseline cannot be captured => fail.
- Scope is unbounded => fail.
- Guard regression after change => fail unless reverted.
- Metric/guard script changed without explicit scope and measurement-integrity note => fail.
- Campaign iteration rejected without rollback evidence or unresolved-risk note => fail.
- Claimed improvement below
min_deltawithout explicit inconclusive status => fail. - Result artifact validator failure => fail.
- Result artifact missing => fail.
Quality Gates
Required:
tests: guard command or impacted tests.review: metric comparison, rollback decision, relevant campaign ledger,git diff --check.artifact: shared validator confirms comparison, experiments JSONL, gate logs, result JSON shape.
Recommended:
lint,format,types: run for any code edits.
Calibration Hooks
On metric/guard-policy change, update calibration:
- behavioral cases: baseline missing, guard regression, noisy metric overclaim, campaign rollback evidence, below-threshold improvement, artifact validator bypass
- benchmark patterns:
optimize
Output Contract
Use shared gate schema from ../../shared/quality-gates.md.
Minimum artifact payload template: result-template.json.
Gives 0 of the 12 instructions most performance cost skills give in ~1.5k tokens
Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07
- keep skill files under 500 linesin 82 of 803, across 16 files
- use imperative form in instructionsin 80 of 803, across 9 files
- draft assertions while test runs are in progressin 75 of 803, across 9 files
- create two to three realistic test promptsin 74 of 803, across 9 files
- write skill descriptions to be pushyin 72 of 803, across 7 files
- save test cases to evals jsonin 72 of 803, across 6 files
- ask questions about edge cases and input formatsin 72 of 803, across 7 files
- save timing data immediately when runs completein 70 of 803, across 5 files
- include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
- launch all test runs in a single turnin 69 of 803, across 3 files
- capture intent before writing a skillin 67 of 803, across 1 file
- import directly instead of barrel filesin 52 of 803, across 15 files
Said here and by no other author read
- capture metric and guard baselines before editing
- write a hypothesis document before iterating
- write a pre-edit patch file each iteration
- re-measure metric and guard commands after editing
- compare baseline and after results in a markdown file
- append iteration results to a jsonl log
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.