Codex autoresearch
Skill TheGreenCedar/codex-autoresearch/plugins/codex-autoresearch/skills/codex-autoresearch
A codex plugin for running optimization loops inside a codebase. It is useful when you have a measurable target and many possible changes to try: test runtime, build speed, bundle size, model loss, Lighthouse scores, memory use, query latency, or any other metric you can print from a script.
npx -y skills add TheGreenCedar/codex-autoresearch --skill codex-autoresearchAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Run or resume a measured improvement loop in a local project. Use for benchmark-driven optimization, qualitative quality-gap research, packet logging, dashboard readouts, recovery, and review-branch finalization backed by autoresearch session files.
SKILL.md
11.3 KB, as published. Nobody here has run it
Codex Autoresearch
Turn an improvement request into a measured, resumable loop. Report the metric, decision, evidence, next action, and real publication state. Do not replace them with a generic claim that the project is "better."
setup -> doctor -> next -> log -> state -> finalize-preview
Use this as the only Codex-facing Autoresearch skill. Do not route to retired subskills, slash commands, or MCP surfaces.
Establish the working truth
- Identify the repository or child package that owns the work.
- Run
git status --short --branch; preserve unrelated changes. - When changing Autoresearch itself, use the checkout in this repository:
- wrapper root:
node plugins/codex-autoresearch/scripts/autoresearch.mjs ... - package root:
node scripts/autoresearch.mjs ...
- wrapper root:
- Treat source and installed-plugin behavior as different until their version and built-entrypoint fingerprint match.
Start or resume
For a new session:
- Get the goal, benchmark, primary metric, direction, correctness checks, editable scope, and any real budget.
- Use
prompt-planorsetup-planwhen one of those is unclear. Both are read-only. - Run
setuponly after the contract is clear enough to create files. - Configure
commitPathsbefore a keep may commit source changes. - Run
doctor --cwd <project> --check-benchmark --explainbefore trusting the first packet. - Record the baseline with
next, thenlog --from-last --status measure.
For an existing session:
- Read loop operations.
- Read
autoresearch.md,autoresearch.jsonl,autoresearch.ideas.md, and the activeautoresearch.research/<slug>/folder when present. - Run
state --report,recommend-next --compact --operator-checklist, anddoctor --explain. These defaults are bounded and share oneresolvedDecision; usestate --json-fullordoctor --json-fullonly for complete machine diagnostics. - Keep
goalFrame.authoritativeGoalauthoritative unless the user deliberately replaces it. If a new request would change the benchmark, metric, edit scope, or final claim, treat it as a possible replacement and resolve that choice before packet work. - Follow the printed blocker or command. If the CLI, report, and dashboard disagree, stop mutation and diagnose the shared state.
Happy path from the package root:
node scripts/autoresearch.mjs setup --cwd <project> --name "<session>" --metric-name <metric> --direction lower --benchmark-command "<command>" --checks-command "<checks>"
node scripts/autoresearch.mjs config --cwd <project> --commit-paths "<editable-paths>"
node scripts/autoresearch.mjs doctor --cwd <project> --check-benchmark --explain
node scripts/autoresearch.mjs next --cwd <project>
node scripts/autoresearch.mjs log --cwd <project> --from-last --status measure --description "Baseline measurement"
node scripts/autoresearch.mjs state --cwd <project> --report
After the baseline, implement one bounded hypothesis inside the configured paths, then run and log one packet:
node scripts/autoresearch.mjs next --cwd <project>
node scripts/autoresearch.mjs log --cwd <project> --from-last --status keep --description "<what changed>" --asi-json-file <path>
node scripts/autoresearch.mjs state --cwd <project> --report
The ASI file must contain the real hypothesis, evidence, rollback reason when rejected, and next action. Use discard, crash, or checks_failed instead of keep when the evidence requires it. Run finalize-preview only when canonical state routes to finalization.
Run one packet at a time
Use next for a reusable packet. Use benchmark-inspect for a bounded diagnostic probe; the old run name fails fast with that migration and is scheduled for removal after 2026-10-01.
After next:
- Inspect the metric, checks, artifacts, diff, and Git state.
- Log with
--from-last; do not copy parsed metrics back into the command. - Add a structured experiment note (ASI) with the hypothesis, evidence, rollback reason for rejected work, and useful next action. Use
--asi-json-file <path>when inline JSON would be fragile in the current shell. - Read the returned continuation before doing anything else.
When accepted work was committed outside Autoresearch, verify the commit and log the keep with --commit <hash> so finalization retains real commit evidence.
| Status | Use it for |
|---|---|
measure | Baselines, no-change checks, environment probes, and diagnostics. Never stage, commit, revert, or finalize it. |
keep | A finite primary metric, passing required checks, and a change worth preserving inside safe Git scope. |
discard | A finite metric and a change not worth keeping; logging may clean the configured or explicit experiment paths. |
crash | A benchmark that failed before usable metric evidence existed. Do not invent a sentinel value; logging may clean the configured or explicit experiment paths. |
checks_failed | A metric exists, but the required correctness proof failed; logging may clean the configured or explicit experiment paths. |
Obey these brakes:
- Keep packet processes on the default minimal environment. Use
--packet-env-mode inheritonly when the benchmark genuinely needs the caller's full environment. - Treat
termination_failedas a hard stop. Preserve partial packet evidence, verify the reported PID and descendants are absent, then remove only the retained progress marker before anothernext. - Treat typed
process_lifecycleblockers as process truth: verify absence before recording a later terminal row. Never infer active residue from historical prose, and never repair a malformed lifecycle row by weakening validation. - Keep a configured working directory inside
--cwd; require the user's explicit intent before passing--allow-outside-workdir. - Ordinary
doctorruns must not refresh remote catalogs. Usedoctor --revalidate-catalogonly for an explicit public-HTTPS provenance check; internal catalogs stay local files. - Continue the active session when
continuation.shouldContinue=true, but run a packet only whenloopContract.canRunNextPacket=true; do not report completion whencontinuation.forbidFinalAnswer=true. - Let blockers, budget stops, segment changes, and finalization outrank another packet.
- Treat
benchmark-lintas a parser check, not proof that the benchmark represents the product. - Require the checks implied by the claim: accuracy, behavior, accessibility, safety, data integrity, or performance.
- Keep
review_requiredresults provisional until the structured note records the review. - Treat benchmark-keyed fixes, static citations, scorer edits, and row-specific detectors as diagnostic until repeat, holdout, breadth, or a promotion gate supports the broader claim.
Use loop operations for partial results, failed checks, ledger repair, budgets, Git scope, and segment changes. Use dashboard and trust for fixed controls, runtime drift, protected paths, redaction, and promotion claims.
Research broad or qualitative work
Use a quality-gap loop for docs, UX, product study, architecture, or research:
node scripts/autoresearch.mjs research-start --cwd <project> --slug <slug> --goal "<goal>"
Keep dated claims in sources.md, judgment in synthesis.md, and accepted work in quality-gaps.md. Preview additions with gap-candidates, then log implementation or rejection with ASI.
Each gap has a stable ID. A checked Markdown box is only a provisional candidate; record the evidence-bearing outcome with gap-decide --gap-id <id> --decision implemented|rejected --evidence <ref> --validation <result>. The append-only decision ledger is the acceptance authority.
If the project already has an executable outcome metric, research-start preserves it as primary and uses quality_gap as secondary acceptance evidence. Treat quality_gap=0 as closure of the accepted checklist for this round only after its decisions are accepted. Read researchIntegrity and its missing-proof warnings before deciding whether the wider question is finished or needs another discovery round.
Read research, lanes, and finalization before fanout, parallel implementation, or review-branch work.
Show the dashboard only when it helps
node scripts/autoresearch.mjs serve --cwd <project>
Verify the server and give the user its http://127.0.0.1:<port>/ URL. Use export for a portable snapshot.
Keep both forms read-only. Run setup, packets, logging, gap work, export, and finalization through the CLI. A static export cannot prove current packet freshness.
Finalize accepted work
- Run
finalize-preview --cwd <project>before branch creation. - For normal finalization, include only accepted, current keeps and exclude session artifacts by default.
- Compare the intended claim with the accepted checks and measurements. If proof is missing, say: "Experimental review branch only: product-grade proof is missing."
- When canonical state reports
current-tree-finalization, treat it as a separate recovery contract: review the entire clean non-session branch diff, exact file set, exclusions, claim evidence, and generated plan, then usefinalize-current-tree --cwd <project> --exclude-session-artifacts. - Ask before creating branches unless the user already approved finalization.
- Verify the branch union, exclusions, summaries, metrics, and checks before handoff.
Report the real runway: preview, approved, branches created, locally verified, pushed or PR, CI, merged, merge verified, then cleanup. Do not collapse those stages or suggest cleanup before the merge is verified.
Keep parent ownership clear
Run codex-goal-brief and inspect top-level canMarkCodexGoalComplete and completionBlocker before the parent calls update_goal(status="complete"). Use --enforce-completion when an invalid completion claim must fail the command. Keep Goal state in Codex; use Autoresearch only for the evidence.
When subagents are explicitly used, give every lane a scope, evidence source, decision, artifact, and test. Scout commands must match lane-runner's strict Git read-only argv allowlist; do not use shell or interpreter escapes. Treat Git porcelain and write-scope checks as best-effort detection, not process/filesystem containment, and use disposable worktrees for implementation lanes. Do not nest subagents or overlap write scopes. Keep the benchmark, packet decision, integration, and final verification in the parent.
Load only the documentation you need
- first run: Start
- normal operation or resume: Operate
- safety or runtime questions: Trust
- review branches: Finish
- symptom lookup: Troubleshooting
- cross-surface disagreement: Control plane
Before claiming plugin work is done, run from plugins/codex-autoresearch:
npm run check
For docs-only work, also inspect the rendered Markdown and command text, then run git diff --check. The package gate checks local Markdown links.