Investigate
Minimal codex-native investigation loop. Use for unknown failures, code debugging, and root-cause narrowing with measurable gates.From its SKILL.md
npx -y skills add Borda/AI-Rig --skill investigateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 24 stars24 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
6.7 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it
Investigate
Diagnosis-first loop for unclear failures: failing tests, tracebacks, regressions, surprising runtime behavior. Produce root-cause claim with evidence, falsification, rejected alternatives before any fix. Use investigate until root cause established; then hand off to develop or code-remediate.
Input Schema
{
"symptom": "required failing command, traceback, runtime bug, CI failure, flaky behavior, or tool anomaly",
"scope": "optional path/module/tool/CI run",
"pace": "fast|full",
"done_when": "one root cause is confirmed or the remaining uncertainty is explicit"
}
Workflow
01: Create run directory
Run python PLUGIN_ROOT/shared/create_run.py --skill investigate once. Retain its single printed path as
<run-directory> and substitute that literal path into every later artifact path and helper argument. Never store or
reuse the path through a shell variable; shell variables do not persist across tool calls.
02: Capture symptom and reproduction context
Write <run-directory>/symptom.md with:
- failing command or observed behavior
- expected behavior
- local vs CI vs external context
- first known bad time or commit, if known
- whether the failure is deterministic, flaky, or unknown
03: Gather signals before forming hypotheses
Run git log --oneline -10 and python --version as separate argv commands. Write their complete outputs to
<run-directory>/recent-commits.txt and <run-directory>/python-version.txt; record either collection failure rather
than treating an empty file as successful evidence.
Inspect python PLUGIN_ROOT/shared/collect_diff.py --help, collect working-tree scope into <run-directory>/baseline; record collection failure, never treat as empty diff.
Add needed tool logs, CI excerpts, tracebacks, config, changed source. Absence of evidence ≠ evidence of absence.
Structural context (optional): when scope names a Python module/symbol, also probe codemap-py once for callers,
coupling, and test impact: python PLUGIN_ROOT/shared/codemap_adapter.py context --category develop --target <qname> --out <run-directory>/codemap-context.json. Per ../../shared/codemap-contract.md, absence/incompatibility is
non-fatal — continue with the signals above. Persist the result once here, before hypothesis ranking; step 05
specialist probes consume <run-directory>/codemap-context.json, never a fresh query.
04: Rank hypotheses in <run-directory>/hypotheses.md
| Rank | Hypothesis | Supporting evidence | Falsification check | Status |
| --- | --- | --- | --- | --- |
Include ≥3 plausible hypotheses unless failing command + code/log directly prove root cause.
05: Orchestrate specialist probes when hypotheses split by domain
Apply ../../shared/specialist-orchestration.md for multi-domain symptoms or useful parallel evidence. Stay single-agent for narrow deterministic failure with one obvious hypothesis.
Write <run-directory>/specialist-probes.md before fan-out: role, hypothesis, context path, expected falsification signal, mode (spawned, substituted, not_triggered).
Recommended probe routing:
qa-specialist: flaky tests, failing assertions, regression reproduction, missing edge-case evidence.cicd-steward: CI-only failure, matrix/cache/permission divergence, release workflow failures.linting-expert: ruff, mypy, pre-commit, tool version or suppression anomalies.security-auditor: auth, secret handling, deserialization, dependency, or permission-related failures.data-steward: data split, leakage, augmentation, DataLoader, or reproducibility anomalies.squeezer: performance regressions, memory/OOM, throughput drops, GPU sync suspicion.scientist: metric instability, paper/method mismatch, experiment validity.web-explorer: volatile dependency or external API behavior.challenger: root-cause claim that would be damaging if wrong.
Each context pack: symptom slice, relevant logs/touched files/environment facts, exact falsification question. Specialists may request context; parent decides widening, consolidates outcomes, owns final root-cause claim.
06: Probe the top hypotheses
Use targeted probes confirming, ruling out, or narrowing one hypothesis at a time.
Each probe must have a clear outcome:
confirmedruled_outinconclusive
Persist probe commands and outputs under <run-directory>/probes/ or inline in <run-directory>/probes.md.
07: Run the anti-rationalization gate
A root-cause claim requires:
- supporting evidence from logs/code/commands
- one falsification check
- at least one rejected alternative
- explicit confidence
Low confidence: continue probing, no fix proposal.
Write <run-directory>/root-cause.md with:
EvidenceFalsificationRejected AlternativesConfidence
08: Run shared quality gates or targeted checks relevant to the failure
Inspect python PLUGIN_ROOT/shared/run_gates.py --help, run full/targeted gates needed to falsify hypotheses.
09: Decide gate result, write result.candidate.json, validate artifacts, and publish .reports/codex/investigate/<timestamp>/result.json
Follow ../../shared/helper-cli-contract.md and authoritative help. Write with INVESTIGATE_METADATA, validate as skill investigate, and promote only the validated candidate.
Fail-Fast Rules
- Missing symptom => fail.
- No evidence collected before hypotheses => fail.
- Root cause stated without falsification check => fail.
- Workaround presented as root cause => fail.
- Missing
root-cause.mdevidence, falsification, rejected alternatives, confidence => fail. - Broad multi-domain symptom without
specialist-probes.mdor an explicit single-agent rationale => fail. - Result artifact validator failure => fail.
- Result artifact missing => fail.
Quality Gates
Required checks:
review: hypothesis table, probe outcomes, rejected alternatives, andgit diff --check.artifact: shared validator confirms investigation artifacts, gate logs, and result JSON shape.
Conditional checks:
tests: failing or confirming reproduction command when available.lint,format,types: only when code/config changes are made as part of a probe.
Calibration Hooks
Update calibration when root-cause routing or workaround rejection changes:
- behavioral cases: symptom-first routing, rejected alternatives, low-confidence probe escalation, artifact validator bypass
- benchmark patterns:
investigate
Output Contract
Use shared gate schema from ../../shared/quality-gates.md.
Minimum artifact payload template: result-template.json.
What ships with it: 1 file
318 B alongside SKILL.md
- result-template.json318 B