agentsclimarketplace

Pr review

Skill momentmaker/kaijutsu/skills/core/pr-review

Adversarial multi-agent pull-request review. Runs claude / codex / antigravity in parallel via `jutsu swarm pr-review`, each with a tailored prompt (claude=architecture+correctness, codex=edge cases, antigravity=cross-file patterns), then synthesizes into a single markdown review with a disagreement table. --full mode adds a Pass-2 round-robin debate where agents critique each other. --strict adds a lie-to-them filter on the synthesis draft. Posts the result as a PR comment when --post-comment is set; edits prior kaijutsu-pr-review comments in place. Use when the user says "review this PR", "code review", "review the diff", "look at PR #N", or invokes /pr-review. Refuses to send the diff unless `.kaijutsu/pr-review.yaml` has `allow-multi-model: true` (first-run prompt persists this). Hard-blocks on a pre-flight secrets scan unless --allow-secrets is passed.From its SKILL.md

Install
npx -y skills add momentmaker/kaijutsu --skill pr-review

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

3 things to look at

  • reads credentialsReads from 2 credential sources: `.env*` and 1 more.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
  • runs commandsInstructs the agent to run 7 commands, including `jutsu swarm pr-review --grant-consent` and 6 more.

SKILL.md

7.7 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it

pr-review

Adversarial multi-agent review. The author already saw the obvious. Your job is to surface what they missed by hearing from THREE independent reviewers — each with a different lens — and distilling the conversation.

This skill wraps jutsu swarm pr-review. The orchestration primitive is in the Go CLI; this skill provides the per-agent prompts, the workflow, and the user-facing wrapper.

Layout (rich)

pr-review/
├── SKILL.md                     (this file — workflow + decision rules)
├── skill.yaml
├── scripts/
│   └── run.sh                   thin wrapper: shells `jutsu swarm pr-review "$@"`
├── prompts/
│   ├── claude.md                "high-level architecture + correctness"
│   ├── codex.md                 "be brutal on edge cases"
│   ├── antigravity.md                "cross-file pattern hunt + consistency"
│   ├── synthesizer.md           cluster, dedupe, render markdown
│   └── debate.md                Pass-2 critique template
├── references/
│   ├── prompt-design.md         why each lens is what it is
│   └── disagreement-rubric.md   how to read the disagreement table
└── runbooks/
    ├── tuning-prompts.md        override prompts per repo
    └── interpreting-output.md   what each section of the comment means

The CLI loads prompts from this skill's prompts/ directory at runtime; if a file is missing it falls back to the built-in default. So authors can override one lens (e.g. add domain-specific rules to antigravity.md) without recompiling jutsu.

How agents collaborate

Three independent passes, one synthesized output:

PassModeWhat runs
1alwaysEach agent reviews independently. N model calls.
2--full onlyEach agent sees others' Pass-1, critiques. N more calls.
SynthalwaysOne agent (default claude) synthesizes into markdown. 1 call.
Filter--strict onlylie-to-them filter trims sycophancy from synthesis. 1 more call.

Cost math (N=3 agents):

  • --quick (default): 4 calls
  • --full: 7 calls
  • --full --strict: 8 calls

--max-cost N warns when the estimate exceeds N USD.

First-run consent

Before the first invocation in a repo, pr-review requires explicit per-repo consent because the diff goes to remote model providers (Anthropic, OpenAI, Google).

Interactive TTY: first invocation prompts and persists the answer. Headless (from inside an agent CLI session, from CI, or piped): the prompt has nothing to read from and the run aborts with a hint. Pre-grant consent with:

jutsu swarm pr-review --grant-consent

This writes allow-multi-model: true to .kaijutsu/pr-review.yaml and exits without running the swarm.

Workflow

When the user says "review this PR", "/pr-review #42", or similar:

  1. Run — invoke via scripts/run.sh "$@" or directly:

    jutsu swarm pr-review --pr 42                 # quick mode, print to stdout
    jutsu swarm pr-review --pr 42 --post-comment  # post to GH after printing
    jutsu swarm pr-review --pr 42 --full          # round-robin debate
    jutsu swarm pr-review --pr 42 --full --strict # debate + lie-to-them filter
    jutsu swarm pr-review --diff-from-branch origin/main  # no PR yet
    
  2. First run only — the CLI prompts:

    This run will send PR diff to N model provider(s). Persist consent to .kaijutsu/pr-review.yaml? [y/N]

    On y it writes allow-multi-model: true to the config. Subsequent runs skip the prompt.

  3. Pre-flight secrets scan — if the diff contains .env*, *.pem, AWS keys, GH PATs, PEM private blocks, etc., the run hard-blocks. Pass --allow-secrets ONLY when the apparent hits are intentional (e.g. test fixtures with placeholder credentials).

  4. Read the output — the markdown comment has three sections:

    • Disagreement table — rows are clustered findings, columns are agents, cells show severity each agent assigned.
    • Synthesis — the synthesizer's prose review.
    • Per-agent stats — collapsed footer with each agent's finding count + cost estimate.

    The 1/N rows in the disagreement table are the conversation-starters. Read those first.

  5. Iterate via --replay — once cached, re-run synthesis without re-calling the model APIs:

    jutsu swarm pr-review --replay <sha>
    

    Useful for tuning the synthesizer prompt.

What the synthesis looks like

## kaijutsu pr-review

**Findings:** 7 total · 2 consensus · 3 contested

### Disagreement Table

| Finding | Severity | Consensus | claude | codex | antigravity |
|---|---|---|---|---|---|
| `auth.go:88` — race in token refresh | blocker | 3/3 | ✓ (blocker) | ✓ (issue) | ✓ (blocker) |
| `cache.go:42` — TTL not honored | issue | 1/3 | — | ✓ (issue) | — |
...

### Synthesis

The change introduces ... (synthesizer's prose) ...

### Disagreements

- **cache.go:42** — codex flagged this as a TTL bug; claude+antigravity did not.
  Worth a closer look.

<details><summary>Per-agent stats</summary>
- **claude** — 4 finding(s) · 14s · est $0.082
- **codex** — 5 finding(s) · 18s · est $0.094
- **antigravity** — 3 finding(s) · 9s · est $0.041
</details>

<!-- kaijutsu-pr-review:run-id=20260505T153022Z sha=abc1234 -->

Configuration

Per-repo .kaijutsu/pr-review.yaml:

allow-multi-model: true              # required gate
agents: [claude, antigravity]             # opt-out of codex if not auth'd
mode: quick                          # default mode for `/pr-review`
max_cost_usd: 1.00
exclude_paths: [pnpm-lock.yaml, vendor/]
synthesizer: claude

Hard rules

  • Never bypass allow-multi-model without explicit user opt-in. The diff goes to 3 model providers — that's a real privacy decision, not a CLI footgun.
  • Never bypass secrets-scan silently. If --allow-secrets is passed, surface the hits in the output so the user knows what they overrode.
  • Adversarial framing is non-negotiable. Per-agent prompts deliberately push for findings, not validation. "Looks good to me" with no evidence is forbidden.
  • The disagreement is the signal. A finding 3/3 agents flag is consensus; one 1/3 agent flags is a conversation-starter. Don't suppress disagreements in the synthesis — call them out.
  • Cite real file:line. Every finding has a path and line range. Made-up locations destroy credibility.
  • Edit-in-place, don't append. Re-running on the same PR edits the prior comment. New SHAs append a "Previous reviews" footer to preserve timeline.

When NOT to use this

  • Tiny one-line diffs — overkill, expensive. Use a single-agent quick review.
  • Diffs with proprietary code under embargo — the diff goes to remote model providers. If your repo's policy doesn't allow that, don't run it.
  • CI gate — Phase 1 is local-only. CI integration arrives later.

What ships with it: 13 files

22.8 KB alongside SKILL.md, 1 of them executable

evals/

prompts/

scripts/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.