agentsclimarketplace

Judge synthesis

Skill event4u-app/agent-config/dist/agent-src/skills/judge-synthesis

Use to consolidate multiple already-run judge verdicts into one report — consensus, conflicts, must-fix/should-fix with per-judge provenance. Consume-only, no opaque score, never auto-gates.From its SKILL.md

Install
npx -y skills add event4u-app/agent-config --skill judge-synthesis

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

7.2 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it

judge-synthesis

The synthesizer across an already-run set of judges. Does not run or re-judge — consumes their verdict blocks and produces one structured report: a side-by-side verdict table, consensus findings (≥2 judges = highest confidence), conflicts (judges disagree), and a synthesized must-fix / should-fix / advisory split with per-finding judge provenance. No opaque single score. Never auto-gates — the human decides.

When to use

  • Two+ judges ran on the same target; their verdict blocks need to become one decision-ready report.
  • A mixed review spanning code judges + artifact/defence judges (PR shipping a roadmap + code + an injection fixture) — no single command consolidates across all seven lenses.
  • /review-changes step 5 wants the canonical synthesis format instead of an ad-hoc merge.

Do NOT use when:

Inputs

The verdict blocks judges already emit (judge name + verdict + findings with severity). Three verdict vocabularies map onto one ordered axis:

Judge familyVerdict vocabulary
code judges (judge-bug-hunter, -code-quality, -security-auditor, -test-coverage, architecture-review-lens)apply / revise / reject
judge-artifact-completenesscomplete / partial / incomplete
judge-injection-defensedefended / partial / breached

Ordered worst→best: reject/incomplete/breached > revise/partial > apply/complete/defended.

Procedure

1. Inspect the inputs and tabulate

Check ≥2 judge verdict blocks are present (only one → stop, nothing to synthesize). Tabulate one row per judge: judge · target · verdict · finding count. Preserve each judge's own verdict word (don't normalise breached into reject); add a severity tier (worst/mid/clean) only as a sort key.

2. Find consensus (highest confidence)

A finding flagged by ≥2 judges (same file:line / dimension / technique) is a consensus finding — highest confidence. List first. By overlap of the finding, never by counting votes for a score.

3. Find conflicts

Two judges with opposite verdicts on one target (one apply, another reject) is a conflict. Surface both verdicts + the disagreement explicitly — never silently resolve by averaging or vote-count. Human adjudicates. Only deterministic rule: for the must-fix list, the most severe verdict wins (one reject → must-fix even if four said apply) — but the conflict is still shown.

4. Synthesize the action split

  • Must-fix — every worst-tier (reject/incomplete/breached) finding + every highest-severity consensus finding.
  • Should-fix — mid-tier (revise/partial) findings.
  • Advisory — single-judge low-severity suggestions.

Each entry carries provenance (which judge(s) raised it). Never merge two judges' findings into one unattributed line.

5. Overall recommendation

One sentence, not a number: block (any worst-tier), revise (any mid-tier, no worst-tier), or proceed (all clean). A recommendation the human acts on — it does not gate.

Validation

  1. Every judge that ran appears exactly once in the table.
  2. Every action entry names its source judge(s).
  3. No single numeric quality score appears anywhere.
  4. Conflicts are shown, not silently resolved.
  5. The recommendation follows the severity rule (worst-tier → block), not a vote.

Output format

Synthesis: <N> judges over <target>
Recommendation: block | revise | proceed

Verdicts:
  judge-bug-hunter          reject    (2 findings)
  judge-security-auditor    apply     (0)
  judge-artifact-completeness  partial (1)
  ...

Consensus (≥2 judges — highest confidence):
  🔴 path:line — <finding> [judge-bug-hunter, judge-code-quality]

Conflicts (judges disagree — human adjudicates):
  <target>: judge-bug-hunter=reject vs judge-security-auditor=apply

Must-fix:
  🔴 <finding> [judge]
Should-fix:
  🟡 <finding> [judge]
Advisory:
  🟢 <finding> [judge]

Required fields (ordered):

  1. Synthesis header + Recommendation — judge count, target, one-word rec
  2. Verdicts — one row per judge, its own verdict word + finding count
  3. Consensus — ≥2-judge findings, with provenance (omit if none)
  4. Conflicts — opposing verdicts on one target (omit if none)
  5. Action split — must-fix / should-fix / advisory, each with provenance

Gotcha

  • No opaque score — the value is the structure (consensus + conflict + provenance), not a rolled-up number that hides which lens objected.
  • Don't normalise verdict words awaybreached and reject sort the same but mean different things; keep each judge's own word.
  • A single reject blocks — even against a majority of apply; but show the conflict, never bury it.
  • Consume, don't dispatch — no judge blocks yet → stop, run the judges first (or hand back to /review-changes).
  • Recommendation ≠ gate — surface it; the human decides.

Do NOT

  • NEVER compute or emit a single numeric quality score
  • NEVER resolve a conflict silently (averaging, vote-count) — surface it
  • NEVER drop a judge's verdict because it disagrees with the majority
  • NEVER run or re-run the judges — this skill consumes their output
  • NEVER auto-gate, auto-reject, or auto-merge on the synthesized recommendation

References

What ships with it: 1 file

1.6 KB alongside SKILL.md

evals/

Keep looking

Skills are one crate of 326,835. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.