Tribunal
When you can't trust yourself with your code base, trust Arbiter.
npx -y skills add arbiterForge/codeArbiter --skill tribunalAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
The deep, rarely-convened whole-codebase audit lane. Routed to when the user invokes {{CMD:tribunal}}. Seven gated phases — cost/model, map, roster dispatch, triage, report, approval+filing, telemetry. Costs on the order of millions of tokens; proceeds only after the user acknowledges the estimate; never a required gate; nothing filed or sent without explicit authorization.
SKILL.md
13.9 KB, as published. Nobody here has run it
tribunal
The deepest, most expensive review codeArbiter offers — convened rarely, on demand, never as a gate. Routed to when the user invokes {{CMD:tribunal}}. Eleven specialist lenses judge the codebase; every finding persists to its own file (plus append-only triage/run logs) under a run dir that survives compaction and disconnects, so the run resumes from disk.
Pre-flight
Read these, or STOP and surface the gap — never guess a command or a path:
{{PROJECT_DIR}}/.codearbiter/tech-stack.md— stack, async model, concurrency primitives, test/lint/secrets commands, and, when documented, the tracker command. Stop if the test/lint/secrets commands are missing; do not guess.{{PROJECT_DIR}}/.codearbiter/CONTEXT.md— thestage:maturity value and domain vocabulary.{{PROJECT_DIR}}/.codearbiter/coding-standards.md— the conventions lenses judge against.{{PROJECT_DIR}}/.codearbiter/security-controls.md— trust boundaries, approved crypto/secret stores; feeds the appsec and secrets lenses. Absent on some repos — proceed without the security lenses' control-file checks if so.- A git repository must be present.
- The reference set under
{{PLUGIN_ROOT}}/skills/tribunal/references/— each is cited at its phase, loaded on demand. Do not preload them.
Phase 0 — Cost, model & resume · gate: STOP
This lane is expensive. Orient and get explicit go-ahead before dispatching anything.
- Resume check. Scan
.codearbiter/reports/for the most recent run dir matching the current scope-slug, any date — never just today's. If none, skip to sizing. If found, check completion: incomplete (noreport-writtenevent in itsrun.jsonl) means either resumable or stale, judged by that run dir's latestrun.jsonltimestamp. A run whoserun.jsonlcarriesrun-abortedis terminal — never offered for resume; a fresh run starts. Younger than 7 days → recover position with the cheap cursor scan inreferences/schemas.md(grep the lastwave-triaged, do not read finding bodies) and offer to resume at the first un-triaged wave instead of restarting; skip the estimate. Older than 7 days → STOP and ask the user to resume anyway or start fresh — the codebase may have drifted under the findings, and stale-tree findings must not silently merge with fresh ones. Complete → start a fresh run. - Abandon. If the user tells the orchestrator to abandon the run, log a
run-abortedevent torun.jsonlbefore stopping. - Cost acknowledgment. Size the job, compute the token band, recommend the model (highest-reasoning available, high effort), and offer the cost-control levers. Present the band plainly; nothing dispatches until the user acknowledges it and confirms the model.
- Establish
RUN_ID=<UTC-date>-<scope-slug>on a fresh run; create.codearbiter/reports/<run-id>/; openrun.jsonl. On resume, reuse the existingRUN_IDas-is — the date is the run's creation date and never changes on resume. - Procedure:
references/cost-and-models.md— load now.
Gate: the user has acknowledged the estimated cost and confirmed the model. An unacknowledged run does not pass.
Phase 1 — Map + judgment overlay · gate: BLOCK
Map before reviewing; the map decides what gets scrutiny.
- Produce the inventory (inline, or on a large repo dispatch the optional cheap mappers per
references/cost-and-models.md): file tree, language breakdown, entry points/routes, core-logic and shared-utility locations, dependency and integration surface. Writeinventory.md. - Apply the judgment overlay in
references/ai-markers.md: risk-rank directories (untrusted input, money, auth, PII, churn = highest), mark trust boundaries, record AI-authorship markers and an iteration-depth estimate. High-marker / high-iteration areas carry a scrutiny boost and a small severity prior. - Choose the active lenses — the full roster minus any whose concern is absent from scope (no migrations → drop the migration lens). Record launched/skipped as
run.jsonlevents. - Choose the wave partition — the default in
references/cost-and-models.md, or a repartition for cause — and record it in therun-startedevent (references/schemas.md); resume reads this recorded partition, never re-derives it.
Gate: inventory.md written with the risk/boundary/marker overlay, and the active-lens set recorded.
Phase 2 — Roster dispatch (dual output: finding files + summary) · gate: BLOCK
Dispatch the active lenses in the wave partition recorded at Phase 1 (default in references/cost-and-models.md) at the concurrency from references/cost-and-models.md (≤5 in flight). Give each agent only its scope slice, on the model/effort from references/cost-and-models.md; the agent itself reads its own mandate (references/lenses/<lens>.md) and the finding contract (references/finding-record.md), and loads neither the other lenses' mandates nor the orchestrator schemas. The orchestrator reads references/finding-record.md to read findings at triage, and consults a lens mandate only to adjudicate that lens's finding.
- Each
tribunal-*agent writes each finding to its own filefindings/<lens>/<finding-id>.jsonthe moment it is found — one file per finding, never a batched write at the end (write contract:references/finding-record.md). - Evidence-or-drop. Every finding cites a concrete
path:lineand the minimal snippet. An absence claim — "no handler", "no teardown", "missing validation" — requires reading the whole unit, never a truncated window. - Specialists never dispatch further subagents. Update each wave's status in
run.jsonlas it flushes. - When a lens's summary returns, record a
lens-completedevent inrun.jsonlwithsurface_seen/findings/modeltaken from the agent's summary, plustokenswhen the orchestrator can observe that lens's spend.{{IF:claude}} - Claude usage receipt. Capture the returned
agentIdfrom each Agent dispatch asagent_idon itslens-launchedevent. If the host returns no usable ID, recordtokens_status: unavailableandtokens_reason: host-result-missingonlens-completed. Before constructing any shell command, require the returned ID to match^[A-Za-z0-9][A-Za-z0-9_-]{0,127}$; on mismatch, recordtokens_status: unavailableandtokens_reason: invalid-agent-id, and do not invoke the helper. Otherwise, after the lens completes, invoke thetribunal-usage.py observe --agent-idmode withpython3 "${CLAUDE_PLUGIN_ROOT}/hooks/tribunal-usage.py" observe --agent-id <validated-agent-id> || python "${CLAUDE_PLUGIN_ROOT}/hooks/tribunal-usage.py" observe --agent-id <validated-agent-id>. The helper resolves only that agent's documentedagent-<agentId>.jsonltranscript and reads only assistant identity and completemessage.usagerecords. Onstatus: observed, copy its integertokensand componenttoken_usage, settokens_status: observed, and copysourceastokens_source. Onstatus: unavailable, omittokens, settokens_status: unavailable, and copyreasonastokens_reason. Usage recovery is best-effort and never blocks a tribunal.{{END}}{{IF:codex}} - Codex usage receipt. If the current host has no subagent dispatch capability and a lens must run inline, record
tokens_status: unavailableandtokens_reason: host-usage-unsupported. If dispatch succeeds but returns no usable thread ID, recordtokens_status: unavailableandtokens_reason: host-result-missing. Otherwise capture the returned agent thread ID on thelens-launchedevent. After that lens completes, resolve the installed plugin root from this routine's own loadedSKILL.mdpath (ordinary shell calls do not inherit a plugin-root environment variable) and runhooks/tribunal-usage.py observe --thread-id <agent-thread-id>. The helper reads only metadata and cumulative token-count events from the exact agent session. Onstatus: observed, copy its integertokensand componenttoken_usage, settokens_status: observed, and copysourceastokens_sourceintolens-completed. Onstatus: unavailable, omittokens, settokens_status: unavailable, and copyreasonastokens_reason; never turn a parser or capability failure into an unexplained omission. Codex session JSONL is explicitly not a stable extension interface, so this recovery remains best-effort and every changed-format path must degrade to a reason, not block the tribunal.{{END}}
Gate: every active lens has flushed its findings/<lens>/ files, and each wave's status is recorded.
Phase 3 — Triage & per-wave planning · gate: BLOCK
Triage per wave from disk as soon as it flushes; do not wait for the whole run.
- Calibrate independently. Set
final_severity/final_confidencefrom the evidence yourself — the lens's values are provisional input; every critical/high carries acounter_argument. - Decide per finding, logged. Each finding gets one decision from the vocabulary, appended as one line to
triage.jsonl. Below the confidence gate after calibration →investigate(medium/low) ordecision-required(critical/high) — never dropped silently. - Plan the wave. Write
plans/phase-<n>.mdfor its kept (keep/combine) work. - Procedure:
references/triage.md— load now.
Gate: every wave's findings triaged into triage.jsonl and a plans/phase-<n>.md written for its kept work.
Phase 4 — Report · gate: BLOCK
Regenerate report.md and manifest.yaml from the two logs per references/report.md — projections, never hand-authored. Task-list-structured (not prose): findings grouped by calibrated severity then type, each with id, path:line, one-line description, remediation shape, triage decision, and a link to its phase plan; decision-required in its own section; a launched/skipped-lens summary; an investigate appendix. Apply {{PLUGIN_ROOT}}/includes/anti-slop-design/ (core + medium-documents) to the prose.
State plainly that critical/high are blocking-severity findings — work that should block shipping the affected code — but that this lane is not itself a gate and blocks nothing.
Gate: report.md regenerated from the logs and presented. No issues created.
Phase 5 — Approval & issue filing · gate: BLOCK
Findings become GitHub issues only on explicit selection and authorization. Silence or ambiguity → file nothing; "looks good" is not authorization.
- Dedup first. Skip findings already carrying an
issue_refintriage.jsonl, then dedup against the tracker — this lane reruns over time and will re-find the same issues. - Default is hand-off. Write and print
issue-commands.sh; execute only on explicit approval, writing eachissue_refback intotriage.jsonl. - Findings file as GitHub issues, never
open-tasks.md— a periodic-review finding must survive PR abandonment. - Procedure:
references/issue-filing.md— load now.
Gate: either issue-commands.sh written and printed, or — on approval — issues filed with the id→result table and issue_ref recorded. Nothing filed without explicit selection; no duplicates against the tracker.
Phase 6 — Telemetry · gate: STOP
Optional, opt-in KPI feedback to refine the skill and the estimator — off by default, sent only on explicit per-run authorization.
- Scrub. The payload is aggregates and per-lens exposure counts only — no code, paths, or finding text; no repo identity unless the user adds
--tag. - Show before send. Write the payload to the run dir and show it in full; state plainly that it posts publicly to the codeArbiter repo. Default: hand the user the ready command; post only on explicit approval.
- Procedure:
references/telemetry.md— load now.
Gate: the payload is shown, and it is either handed to the user as a command or — on approval — posted. No telemetry leaves without per-run authorization.
Hard rules
- MUST NOT proceed past Phase 0 without the user acknowledging the estimated token cost — this lane can cost millions of tokens.
- MUST NOT edit, refactor, format, or commit project code — writes are confined to
.codearbiter/reports/<run-id>/until the filing gate. - MUST NOT act as a required gate or block a merge, commit, or other workflow — critical/high are blocking-severity findings, not a pipeline halt.
- MUST NOT record a finding without a concrete
path:lineand a minimal evidence snippet. - MUST NOT assert an absence without reading the whole relevant unit — partial-window absence claims do not pass.
- MUST NOT let a lens's provisional severity/confidence stand as final — calibrate at triage; every critical/high carries a
counter_argument. - MUST NOT mutate the append-only logs —
manifest.yaml,report.md, andplans/are regenerated from them, never hand-edited. - MUST NOT file an issue below the confidence gate or without explicit selection and authorization; findings file as GitHub issues, never
open-tasks.md. - MUST NOT create a duplicate issue — skip findings carrying an
issue_ref, and dedup against the tracker bydedup_key/title before filing. - MUST NOT author or scaffold an ADR —
decision-requiredfindings file as a discussion issue; ADRs are authored only via{{CMD:adr}}with user attribution. - MUST NOT send telemetry without explicit per-run authorization, and MUST NOT include code, file paths, finding text, or repo identity (absent an explicit
--tag) in the payload — KPI aggregates only. - MUST NOT guess the test, lint, or secrets-scan command — read
tech-stack.mdor STOP. For the tracker: usetech-stack.mdif it documents one; else default togh issue createon a GitHub origin; else STOP. - MUST NOT dispatch a subagent from within a dispatched specialist — only the orchestrator dispatches.