Ca review
Skill arbiterForge/codeArbiter/plugins/ca-pi/skills/ca-review
Review a diff with the reviewer fleet, funneled to one triaged verdict. Targets the current working diff, a path, or an inbound GitHub PR.From its SKILL.md
npx -y skills add arbiterForge/codeArbiter --skill ca-reviewAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
SKILL.md
4.7 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
/ca-review — diff review
Read-only review of a change. Routes to dispatching-parallel-agents: dispatches the reviewer fleet by path matrix, dedupes, then funnels through finding-triage → checkpoint-aggregator to a single verdict. No code is modified.
The change under review does not have to be yours. /ca-review #123 reviews an inbound pull request through the same fleet, the same matrix, and the same triage. That is the point of issue #80: a tool that only reviews the diff you just wrote is a linter for authors, not a gate for a team, and reviewing code you did NOT write is where a governance gate earns its keep.
It is an ARGUMENT, not a second command. The scope resolver already took one, the fleet is scope-agnostic, and every phase downstream operates on a diff regardless of where it came from — so a /ca-review-pr would be a whole public surface (catalog, three host projections, README counts, sidebar) whose only distinguishing feature is where the diff was fetched from.
Flow
-
Resolve scope from
$ARGUMENTS:- empty → the current working diff (unchanged default).
- a path → that path (unchanged).
#<number>, a bare number, or a GitHub PR URL → an INBOUND PR. Fetch its diff withgh pr diff <number>and review that. Ifghis missing or unauthenticated, STOP and say so — do NOT silently fall back to the working diff, which would report a verdict on the wrong change under the PR's name.
For a PR target, resolve the diff ONCE and review that text. Do not re-fetch per reviewer: the fleet runs in parallel, and a PR updated mid-review would otherwise have different reviewers reading different code and a triage that reconciles findings from two versions.
-
Build the unit list by path matrix; each matched reviewer is one read-only unit:
Reviewer Dispatched when scope touches security-reviewerauth, middleware, secrets, deploy/CI, any security-sensitive path auth-crypto-reviewerauthn, crypto, key handling, secrets dependency-reviewerpackage.json, lockfiles, base images, dependency manifestsmigration-reviewerDB migration file add/modify coverage-auditorany source change (test coverage vs. obligations) architecture-drift-reviewercode that may diverge from accepted ADRs in .codearbiter/decisions/ -
Route to
dispatching-parallel-agentswith that unit list (read-only batch — no collision check). It dedupes overlapping findings, then funnels throughfinding-triage(severity + inline[NEEDS-TRIAGE]on out-of-scope items) →checkpoint-aggregator(single verdict). -
Surface the aggregated verdict: findings by severity, file:line, remediation, and the applicable control from
<project-root>/.codearbiter/security-controls.mdfor security findings. -
For a PR target, posting the verdict is a separate, confirmed step. Report locally first; post only on explicit instruction, with
gh pr review <number> --comment --body-file <file>. A review comment on someone else's PR is outward-facing and effectively public the moment it lands — it notifies subscribers and cannot be un-sent. Never--request-changesor--approvefrom here: those carry merge authority, and this command produces a finding list, not a maintainer's decision.
Severity
- CRITICAL — exploitable vuln, secret exposure, banned primitive, data-integrity breach.
- HIGH — significant compliance gap or unsafe pattern.
- MEDIUM — standards deviation or coverage gap.
- LOW — informational or style.
Hard gate
Read-only — MUST NOT modify a file, and MUST NOT check out, merge, or otherwise move the repository to the PR's branch: reviewing an inbound PR means reading its DIFF, not adopting its code, and a checkout would run its content through hooks that trust the working tree. BLOCK on any CRITICAL or HIGH finding on your OWN change: it must be resolved before /ca-pr. On an inbound PR there is nothing local to block — the verdict is the deliverable. MUST NOT consume raw reviewer output — only the finding-triage → checkpoint-aggregator
verdict. MUST NOT resolve a [CONFIRM-NN] surfaced during review by guessing.
When NOT to use
- Opening a PR (reviews dispatch automatically) →
/ca-pr. - A periodic full-codebase sweep →
/ca-checkpoint. - A pre-implementation threat model →
/ca-threat-model. - A question about the code →
/ca-btw.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most review quality skills give in ~1.1k tokens
Counted across 1,048 of the 1,783 authors here whose files we hold, read 2026-08-07
- Ask questions one at a timein 81 of 1048, across 64 files
- Provide a recommended answer for each questionin 73 of 1048, across 50 files
- Explore the codebase instead of asking answerable questionsin 66 of 1048, across 42 files
- Resolve dependencies between decisions one-by-onein 42 of 1048, across 17 files
- Interview the user relentlessly about the planin 38 of 1048, across 13 files
- Order findings by severityin 31 of 1048
- Resolve each branch of the decision treein 27 of 1048, across 5 files
- Run a grilling sessionin 26 of 1048, across 5 files
- Update CONTEXT.md immediately when a term is resolvedin 26 of 1048, across 11 files
- Propose precise canonical terms for vague languagein 25 of 1048, across 7 files
- Create documentation files lazilyin 24 of 1048, across 5 files
- Assign severity to every findingin 24 of 1048
Said here and by no other author read
- build reviewer unit list by path matrix
- funnel findings through triage and aggregator
- surface aggregated verdict by severity
- post PR comments only on explicit instruction
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.