Calibrate
A collection of personal AI coding assistant configurations, specialist agents, and automated workflows optimized for Python and ML open-source development.
npx -y skills add Borda/AI-Rig --skill calibrateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 23 stars23 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Codex-native calibration loop. Use to detect leaks or major gaps across packaged skills and role cards with fixed checks plus behavioral recall, precision, and confidence-accuracy scoring.
SKILL.md
8.7 KB, as published. Nobody here has run it
Calibrate
Run calibration for Codex workflow integrity and behavioral scoring.
Input Schema
{
"scope": "skills|agents|routing|all",
"pace": "fast|full",
"mode": "ab-test|apply",
"require_live_routes": false,
"skip_gate": false,
"done_when": "recall and bias scores emitted; proposals written if mode=apply; gate skipped if skip_gate=true"
}
Workflow
Installed plugin runs use --layout plugin --root <consuming-project>. The runner discovers package assets from its
own file location under runtime/calibration, skills, roles, and shared; --root controls only report output,
Git context, and read-only classification work. It must not fall back to a source checkout or project .codex.
Repository maintainers may use --layout source --root <source-project> to validate the source .codex layout. Do
not mix source agents, sync manifests, or project registration checks into an installed-plugin result.
01: Load calibration task set from ../../runtime/calibration/tasks.json
02: Load behavioral cases from ../../runtime/calibration/behavioral-cases.json
03: Load behavioral observations from ../../runtime/calibration/behavioral-observations.jsonl
- Require
source,run_id,observed_at.source=live-*also needs route; campaign/pair IDs; pair/registered role; actual model/effort; recomputable prompt/task-contract SHA-256; task type/scope; input/cached/output tokens; latency; outcome; tool/check failures; normalized cost; pricing reference. Each complete campaign exactly matches case/role/type/scope signatures inlive-ab-tasks.json; substituted task, fixture, gate, prompt input fails.
04: Inspect ../../runtime/calibration/run.py --help, then run plugin layout against the consuming project
Use --require-live-routes only for the strict-live gate. Default offline scoring remains fixture-backed and makes no
paid model calls.
05: Inspect checks_failed, leaks_found, and behavioral
06: Review behavioral metrics:
recall: expected IDs recovered from known cases.precision: reported IDs matching expected IDs.confidence_accuracy:1 - mean(abs(confidence - per-case F1)).mean_overconfidence: mean positive confidence bias over per-case F1.gate_metrics_raw: unrounded pass/fail values.by_source: recall, precision, confidence calibration by source.observation_freshness: latestobserved_at, missing timestamps, live/fixture counts.live_route_acceptance: matched baseline/candidate classification and isolated tool-use quality, normalized token-efficiency proxy, evidence sufficiency per configured route; not monetary pricing evidence.
07: Classify gaps as blocking or non-blocking
08: Emit measured recommendations for what should be fixed or improved next
- Failed checks/leaks first.
- Behavioral recommendations name metric gap/affected cases when available.
- Separate fixture-only caveats from live-quality claims.
09: Write skill artifacts to .reports/codex/calibrate/<timestamp>/; preserve runner evidence under .reports/codex/calibration/<timestamp>/
10: Write the validated skill-level artifact when this skill wraps the runner
Follow ../../shared/helper-cli-contract.md/authoritative help. Gate intent: ruff lint/format calibration+skills, explicit no-typed-target reason, calibration tests, clean diff. Write CALIBRATE_METADATA, validate calibrate, promote only validated candidate.
Native Contract Checks
Verify configured native surface, not only runner internals.
Skill checks:
- configured skill file exists; frontmatter has unindented
---,name:,description:; required sections exist; artifact path.reports/codex/<skill>/; examples includestatus,checks_run,checks_failed,findings,confidence,artifact_path; no external runner-only metadata/cache. - CLI checks find every local shebang Python/shell entry point in calibration, shared helpers, code-review, offline harness; each executable, fixed-help-roster registered, authoritative
--help. - every skill references
helper-cli-contract.md, not complete local CLI invocations. - source layout compares
../../runtime/calibration/behavioral-cases.jsonversion toHEAD: dirty tree same or exactly one commit-relative version step; installed plugin layout records the packaged fixture as immutable.
Role checks:
- installed layout requires every packaged
roles/<role>/ROLE.md; source layout requires each configured source agent. - role-card frontmatter contains role ID, namespaced name, active model, reasoning effort, approval policy, sandbox, and fallback modes; package-manifest skill/role rosters contain every calibrated target.
- default, review parent, runtime, research, curation, adversarial use
gpt-5.6-terra; delegation/docs/CI-CD/web/OSS/static analysis usegpt-5.6-luna; only security/solution architecture usegpt-5.6-sol. - Luna/high is explicit human override for bounded simpler roles; preserve strict quality/cost failure, reject undocumented expansion.
- every role defaults
high;xhigh/maxexplicit task escalation.model_reasoning_effortfollows agent-effort-policy: allhigh,xhigh/maxtask overrides. - high-stakes roles use high-capability tier; bounded support may lower-cost tier. No deprecated model string in active config/TOML.
- role has clear trigger/skip/not-for boundaries, evidence ownership, execution constraints, handover, and confidence contracts; sensitive roles retain sandbox, especially read-only security audit; packaged roles require no external runtime path variable.
Usage Notes
- After meaningful agent/skill instruction change, confirm routing/output match stack.
leaks_foundprimary drift;checks_failedmechanical gate.- Behavioral metrics measure supplied observations only.
fixture-selftestvalidates scoring; live Codex quality requires replacing/appending live-prompt observations. - Missing route coverage is
insufficient-evidence, never acceptance;require_live_routes=trueexits nonzero. - Compare thresholds with
gate_metrics_raw, not rounded display. - Paid paired campaigns:
../../runtime/calibration/run_live_ab.py; plans by default, executes only--confirm-paid-run=chatgpt-subscription, verified local ChatGPT subscription login, no API key env, noCI/GITHUB_ACTIONS. - Each live task names a canonical role. Plugin layout prepends the exact packaged role card to both prompts; source layout preserves project-instruction plus source-agent prompt construction. Tool pairs can accept a candidate passing an executable gate when the successfully invoked baseline fails; infrastructure timeout is never a candidate win.
- Sol critical-only unless paired quality exceeds Terra configured minimum; tie retains Terra.
- Do not claim currency savings from
normalized-token-v1; need dated authoritative model-specific price. - Fixture
versionis committed-history marker: comparegit show HEAD:<path>; dirty tree stays committed or one-next version until commit. - Missing registration/pattern mismatch: minimal config fix then rerun before widening.
Fail-Fast Rules
- Missing calibration files => fail.
- Missing configured skill or role file => fail.
- Native skill/role contract mismatch => fail unless result waives.
- Runtime leakage in native skill or role files => fail.
- Behavioral gate below threshold => fail.
- Result artifact missing => fail.
- Behavioral case-set version >1 step from committed version => fail.
require_live_routes=truewith incomplete route pairs => fail.- Live row without the strict paired execution schema => fail.
Quality Gates
Required checks:
calibration:../../runtime/calibration/run.py --layout plugin --root <consuming-project>.behavioral-version-policy: compare case-set version toHEAD; avoid meaningless dirty-tree gaps.review: inspect failed patterns, leaks, behavioral gaps, stale fixtures before recommendations.
Conditional checks:
tests: run focused tests when calibration code changes.format: validate JSON and shell syntax when calibration fixtures change.
Calibration Hooks
When calibration expectations change, update together:
../../runtime/calibration/benchmarks.json../../runtime/calibration/behavioral-cases.json../../runtime/calibration/behavioral-observations.jsonl../../runtime/calibration/run.py../../runtime/calibration/live-route-policy.json../../runtime/calibration/live-ab-tasks.json../../runtime/calibration/run_live_ab.py
Output Contract
Use shared gate schema from ../../shared/quality-gates.md.
Minimum artifact payload template: result-template.json.