agentsclimarketplace

Optimize

Skill Borda/AI-Rig/plugins/codex-rig/skills/optimize

A collection of personal AI coding assistant configurations, specialist agents, and automated workflows optimized for Python and ML open-source development.

Install
npx -y skills add Borda/AI-Rig --skill optimize

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 23 stars23 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Minimal codex-native optimization loop. Use for metric-driven improvements with guardrails and measurable gates.

SKILL.md

6.7 KB, ~1.5k tokens by cl100k_base, as published. Nobody here has run it

Optimize

Metric-driven optimization with explicit guards, rollback criteria, experiment log.

Input Schema

{
  "goal": "required measurable improvement objective",
  "mode": "single|campaign",
  "metric_cmd": "required command that emits or validates the target metric",
  "metric_direction": "higher|lower",
  "guard_cmd": "required command that must continue to pass",
  "max_iterations": "optional integer, default 1",
  "min_delta": "optional practical significance threshold",
  "scope_files": [
    "paths the optimization may edit"
  ],
  "done_when": "metric improves without guard regression"
}

Workflow

01: Create run directory

Run python PLUGIN_ROOT/shared/create_run.py --skill optimize once. Retain its single printed path as <run-directory> and substitute that literal path into every later artifact path and helper argument. Never store or reuse the path through a shell variable; shell variables do not persist across tool calls.

02: Validate metric and guard commands

Require:

  • Repeatable metric_cmd producing comparable value or pass/fail.
  • Known metric_direction.
  • guard_cmd fails on unacceptable regressions.
  • Bounded scope_files.
  • Explicit, bounded max_iterations for campaign.
  • Protect files/scripts used by metric_cmd/guard_cmd unless user explicitly scopes them and accepts measurement-integrity risk.

Dry-run both before edit:

Execute the configured metric_cmd and guard_cmd separately with the host-native command runner. Write complete combined output to <run-directory>/metric-baseline.txt and <run-directory>/guard-baseline.txt; retain both exit codes and stop before editing if either command cannot run.

03: Record baseline and hypothesis

Write <run-directory>/hypothesis.md:

  • metric to improve
  • expected mechanism
  • files allowed to change
  • guard risk
  • rollback condition

For campaign, noisy metrics, GPU/ML performance, or correctness-sensitive code, apply ../../shared/specialist-orchestration.md. Write <run-directory>/specialist-optimization-plan.md with narrow context packs for:

  • squeezer: profiling mechanism, bottleneck hypothesis, measurement plan.
  • qa-specialist: guard coverage and regression risk.
  • data-steward: data pipeline or reproducibility impact.
  • scientist: metric validity, ablation design, statistical noise.
  • challenger: overfitting to the metric or weakening guard checks.

No fan-out for one small measured change with stable metric/guard. Never let specialist change metric/guard scripts unless explicitly in scope_files and measurement-integrity risk recorded.

Structural context (optional): when scope_files resolves to a Python module/symbol, also probe codemap-py once for callers, coupling, and test impact before the first iteration: python PLUGIN_ROOT/shared/codemap_adapter.py context --category develop --target <qname> --out <run-directory>/codemap-context.json. Per ../../shared/codemap-contract.md, absence/incompatibility is non-fatal — continue with the hypothesis above. Persist the result once here, before step 04 applies any change; any triggered specialist consumes <run-directory>/codemap-context.json, never a fresh query.

Initialize machine-readable iteration log:

Create an empty <run-directory>/experiments.jsonl with the filesystem tool before the first iteration.

04: Apply one minimal optimization change per iteration

One independent hypothesis per iteration. Do not optimize unmeasured paths. Before each, write <run-directory>/iteration-<n>-before.patch with scoped-file diff. If iteration fails and only its patch is present, revert with git apply -R against iteration diff; otherwise fail run when clean reversal cannot be proven. Never use git reset --hard.

05: Re-measure

Re-run the same retained metric_cmd and guard_cmd separately with the host-native command runner. Write complete combined output to <run-directory>/metric-after.txt and <run-directory>/guard-after.txt; retain both exit codes.

06: Compare baseline and after results in <run-directory>/comparison.md

Required fields:

  • baseline value
  • after value
  • delta
  • guard status
  • confidence
  • noise caveats

Append one JSON object/iteration to <run-directory>/experiments.jsonl:

{
  "iteration": 1,
  "hypothesis": "one-line mechanism",
  "metric_before": 0.0,
  "metric_after": 0.0,
  "delta": 0.0,
  "guard": "pass|fail",
  "decision": "kept|reverted|inconclusive|failed",
  "rollback_evidence": "path or reason"
}

07: Decide keep/revert

  • Keep only with intended metric movement and passing guards.
  • With min_delta, keep only if delta meets/exceeds practical-significance threshold.
  • Revert or fail on guard regression.
  • For noisy measurement, repeat or mark inconclusive.
  • In campaign, stop at first kept result unless user asked continued exploration; otherwise continue only while max_iterations remains and each rejected iteration has rollback evidence.

08: Run shared quality gates

Inspect python PLUGIN_ROOT/shared/run_gates.py --help. Tests runs configured test or guard command; give real commands or explicit reasons for other gates.

09: Write and validate the mandatory result artifact

Follow ../../shared/helper-cli-contract.md and authoritative help. Write OPTIMIZE_METADATA, validate as optimize, promote only validated candidate.

Fail-Fast Rules

  1. Missing metric or guard command => fail.
  2. Baseline cannot be captured => fail.
  3. Scope is unbounded => fail.
  4. Guard regression after change => fail unless reverted.
  5. Metric/guard script changed without explicit scope and measurement-integrity note => fail.
  6. Campaign iteration rejected without rollback evidence or unresolved-risk note => fail.
  7. Claimed improvement below min_delta without explicit inconclusive status => fail.
  8. Result artifact validator failure => fail.
  9. Result artifact missing => fail.

Quality Gates

Required:

  • tests: guard command or impacted tests.
  • review: metric comparison, rollback decision, relevant campaign ledger, git diff --check.
  • artifact: shared validator confirms comparison, experiments JSONL, gate logs, result JSON shape.

Recommended:

  • lint, format, types: run for any code edits.

Calibration Hooks

On metric/guard-policy change, update calibration:

  • behavioral cases: baseline missing, guard regression, noisy metric overclaim, campaign rollback evidence, below-threshold improvement, artifact validator bypass
  • benchmark patterns: optimize

Output Contract

Use shared gate schema from ../../shared/quality-gates.md.

Minimum artifact payload template: result-template.json.

Gives 0 of the 12 instructions most performance cost skills give in ~1.5k tokens

Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07

  • keep skill files under 500 linesin 82 of 803, across 16 files
  • use imperative form in instructionsin 80 of 803, across 9 files
  • draft assertions while test runs are in progressin 75 of 803, across 9 files
  • create two to three realistic test promptsin 74 of 803, across 9 files
  • write skill descriptions to be pushyin 72 of 803, across 7 files
  • save test cases to evals jsonin 72 of 803, across 6 files
  • ask questions about edge cases and input formatsin 72 of 803, across 7 files
  • save timing data immediately when runs completein 70 of 803, across 5 files
  • include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
  • launch all test runs in a single turnin 69 of 803, across 3 files
  • capture intent before writing a skillin 67 of 803, across 1 file
  • import directly instead of barrel filesin 52 of 803, across 15 files

Said here and by no other author read

  • capture metric and guard baselines before editing
  • write a hypothesis document before iterating
  • write a pre-edit patch file each iteration
  • re-measure metric and guard commands after editing
  • compare baseline and after results in a markdown file
  • append iteration results to a jsonl log

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.