Judge synthesis
Skill event4u-app/agent-config/dist/agent-src/skills/judge-synthesis
Use to consolidate multiple already-run judge verdicts into one report — consensus, conflicts, must-fix/should-fix with per-judge provenance. Consume-only, no opaque score, never auto-gates.From its SKILL.md
npx -y skills add event4u-app/agent-config --skill judge-synthesisAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 7 stars7 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.2 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it
judge-synthesis
The synthesizer across an already-run set of judges. Does not run or re-judge — consumes their verdict blocks and produces one structured report: a side-by-side verdict table, consensus findings (≥2 judges = highest confidence), conflicts (judges disagree), and a synthesized must-fix / should-fix / advisory split with per-finding judge provenance. No opaque single score. Never auto-gates — the human decides.
When to use
- Two+ judges ran on the same target; their verdict blocks need to become one decision-ready report.
- A mixed review spanning code judges + artifact/defence judges (PR shipping a roadmap + code + an injection fixture) — no single command consolidates across all seven lenses.
/review-changesstep 5 wants the canonical synthesis format instead of an ad-hoc merge.
Do NOT use when:
- Only one judge ran — nothing to synthesize; surface that block directly.
- You need to produce a verdict — that is the individual judge's job (
judge-bug-hunter,judge-artifact-completeness,judge-injection-defense). - You are tempted to compute a single quality number — forbidden (Do NOT).
Inputs
The verdict blocks judges already emit (judge name + verdict + findings with severity). Three verdict vocabularies map onto one ordered axis:
| Judge family | Verdict vocabulary |
|---|---|
code judges (judge-bug-hunter, -code-quality, -security-auditor, -test-coverage, architecture-review-lens) | apply / revise / reject |
judge-artifact-completeness | complete / partial / incomplete |
judge-injection-defense | defended / partial / breached |
Ordered worst→best: reject/incomplete/breached > revise/partial > apply/complete/defended.
Procedure
1. Inspect the inputs and tabulate
Check ≥2 judge verdict blocks are present (only one → stop, nothing to synthesize). Tabulate one row per judge: judge · target · verdict · finding count. Preserve each judge's own verdict word (don't normalise breached into reject); add a severity tier (worst/mid/clean) only as a sort key.
2. Find consensus (highest confidence)
A finding flagged by ≥2 judges (same file:line / dimension / technique) is a consensus finding — highest confidence. List first. By overlap of the finding, never by counting votes for a score.
3. Find conflicts
Two judges with opposite verdicts on one target (one apply, another reject) is a conflict. Surface both verdicts + the disagreement explicitly — never silently resolve by averaging or vote-count. Human adjudicates. Only deterministic rule: for the must-fix list, the most severe verdict wins (one reject → must-fix even if four said apply) — but the conflict is still shown.
4. Synthesize the action split
- Must-fix — every worst-tier (
reject/incomplete/breached) finding + every highest-severity consensus finding. - Should-fix — mid-tier (
revise/partial) findings. - Advisory — single-judge low-severity suggestions.
Each entry carries provenance (which judge(s) raised it). Never merge two judges' findings into one unattributed line.
5. Overall recommendation
One sentence, not a number: block (any worst-tier), revise (any mid-tier, no worst-tier), or proceed (all clean). A recommendation the human acts on — it does not gate.
Validation
- Every judge that ran appears exactly once in the table.
- Every action entry names its source judge(s).
- No single numeric quality score appears anywhere.
- Conflicts are shown, not silently resolved.
- The recommendation follows the severity rule (worst-tier → block), not a vote.
Output format
Synthesis: <N> judges over <target>
Recommendation: block | revise | proceed
Verdicts:
judge-bug-hunter reject (2 findings)
judge-security-auditor apply (0)
judge-artifact-completeness partial (1)
...
Consensus (≥2 judges — highest confidence):
🔴 path:line — <finding> [judge-bug-hunter, judge-code-quality]
Conflicts (judges disagree — human adjudicates):
<target>: judge-bug-hunter=reject vs judge-security-auditor=apply
Must-fix:
🔴 <finding> [judge]
Should-fix:
🟡 <finding> [judge]
Advisory:
🟢 <finding> [judge]
Required fields (ordered):
- Synthesis header + Recommendation — judge count, target, one-word rec
- Verdicts — one row per judge, its own verdict word + finding count
- Consensus — ≥2-judge findings, with provenance (omit if none)
- Conflicts — opposing verdicts on one target (omit if none)
- Action split — must-fix / should-fix / advisory, each with provenance
Gotcha
- No opaque score — the value is the structure (consensus + conflict + provenance), not a rolled-up number that hides which lens objected.
- Don't normalise verdict words away —
breachedandrejectsort the same but mean different things; keep each judge's own word. - A single reject blocks — even against a majority of
apply; but show the conflict, never bury it. - Consume, don't dispatch — no judge blocks yet → stop, run the judges first (or hand back to
/review-changes). - Recommendation ≠ gate — surface it; the human decides.
Do NOT
- NEVER compute or emit a single numeric quality score
- NEVER resolve a conflict silently (averaging, vote-count) — surface it
- NEVER drop a judge's verdict because it disagrees with the majority
- NEVER run or re-run the judges — this skill consumes their output
- NEVER auto-gate, auto-reject, or auto-merge on the synthesized recommendation
References
- Code judges:
judge-bug-hunter,judge-code-quality,judge-security-auditor,judge-test-coverage,architecture-review-lens. - Artifact / defence judges:
judge-artifact-completeness,judge-injection-defense. - Dispatchers that feed this:
/review-changes(5 code judges),subagent-orchestration(parallel judge fan-out). - Cross-model review families (disambiguation):
ai-council(independent breadth),/team(collaborative repo-access depth) — this skill consolidates in-session same-weights judges. - LLM-as-a-Judge — Zheng et al. (2023), arxiv.org/abs/2306.05685; the consolidation layer over the specialized-judge pattern, with consensus/conflict surfacing instead of a single aggregate score.
What ships with it: 1 file
1.6 KB alongside SKILL.md
evals/
- triggers.json1.6 KB