Second opinion bench
Measure whether a second, independent reviewer model catches complementary defects in the second-opinion-bench synthetic corpus.From its SKILL.md
npx -y skills add runsagents/second-opinion-benchAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
3.6 KB, 754 tokens by cl100k_base, as published. Nobody here has run it
Second Opinion Bench
Use this skill when asked to produce benchmark verdicts with the current model or to grade verdicts already produced elsewhere.
Integrity rules
- The harness is model-agnostic and local. Never add a live model call to
src/; verdict generation happens in the host agent. - During review, read only
fixtures/diffs/*/change.diff. Do not read anyground-truth.json, prior verdict, result, or example report. - Review every fixture independently. Do not infer that each diff must contain exactly one issue, even though the current corpus does.
- Freeze the raw finding summaries before opening ground truth. After freezing, map only an already-recorded finding to a planted defect ID when the semantics clearly match. Never create, expand, or delete a finding because of ground truth.
- Give every unmatched finding a stable reviewer-proposed ID namespaced to the fixture. The grader must be allowed to count it as a false positive.
- Never describe a simulated, hand-authored, incomplete, or ground-truth-exposed run as measured model performance.
Produce verdicts with this model
Choose a reviewer label that identifies the actual model/version and material review configuration. Create reviews/<reviewer>/.
For each fixture directory, in isolation:
- Read
change.diffonly. - Look for correctness, edge-case, security, monetary, concurrency, and behavior-contract defects. Record only defects supported by the diff.
- Freeze a list of concise summaries. An empty list is allowed.
- Only now read that fixture’s
ground-truth.jsonto map a frozen summary to an exact planted ID. Preserve unmatched summaries under reviewer-proposed IDs. - Save
reviews/<reviewer>/<fixtureId>.jsonconforming toschemas/verdict.schema.json:
{
"reviewer": "provider-model-version-prompt-v1",
"fixtureId": "edge-empty-average",
"foundDefects": [
{
"id": "edge-empty-average/reduce-without-seed",
"summary": "Empty input calls reduce without an initial value."
}
]
}
Write one verdict for every fixture, including "foundDefects": [] when no defect was found. Do not copy the demonstration verdicts.
To test a different model, run this protocol again in a fresh agent context using that model. Do not expose the new reviewer to another reviewer’s files or results until its verdicts are complete.
Grade sealed verdicts
After all reviewers have sealed their files, run:
node src/run-bench.mjs \
--reviews reviews \
--output bench-result.json \
--report bench-report.md
The command reads JSON files only; it does not contact a model or provider. Treat any coverage warning as an incomplete run, not as a zero-finding review.
Report honestly
State:
- reviewer labels and enough model/configuration detail to reproduce them;
- fixture coverage for every reviewer;
- overall and per-class recall with numerators and denominators;
- false-positive counts;
- pair union recall;
- both directional uplifts: what B adds after A and what A adds after B;
- whether any verdict was simulated, hand-authored, post-edited, or exposed to ground truth before findings were frozen;
- the limitations that fixtures are synthetic and the grader matches IDs, not semantics.
Do not say “model B is better” from this corpus alone. A defensible statement is narrower: “On this 12-defect synthetic corpus, under the recorded setup, B added N planted defects that A missed.”
What ships with it: 67 files
75.8 KB alongside SKILL.md, 7 of them executable
examples/
- bench-report.md1.3 KB
- bench-result.json7.6 KB
- reviews/simulated-reviewer-a/concurrency-idempotency-race.json260 B
- reviews/simulated-reviewer-a/concurrency-lost-update.json126 B
- reviews/simulated-reviewer-a/edge-empty-average.json223 B
- reviews/simulated-reviewer-a/edge-missing-header.json122 B
- reviews/simulated-reviewer-a/logic-inverted-entitlement.json243 B
- reviews/simulated-reviewer-a/logic-pagination-boundary.json236 B
- reviews/simulated-reviewer-a/money-float-fee.json256 B
- reviews/simulated-reviewer-a/money-line-rounding.json122 B
- reviews/simulated-reviewer-a/security-path-traversal.json126 B
- reviews/simulated-reviewer-a/security-unverified-token.json269 B
- reviews/simulated-reviewer-a/silent-default-flip.json248 B
- reviews/simulated-reviewer-a/silent-error-swallow.json123 B
- reviews/simulated-reviewer-b/concurrency-idempotency-race.json262 B
- reviews/simulated-reviewer-b/concurrency-lost-update.json264 B
- reviews/simulated-reviewer-b/edge-empty-average.json121 B
- reviews/simulated-reviewer-b/edge-missing-header.json122 B
- reviews/simulated-reviewer-b/logic-inverted-entitlement.json129 B
- reviews/simulated-reviewer-b/logic-pagination-boundary.json294 B
- reviews/simulated-reviewer-b/money-float-fee.json220 B
- reviews/simulated-reviewer-b/money-line-rounding.json254 B
- reviews/simulated-reviewer-b/security-path-traversal.json261 B
- reviews/simulated-reviewer-b/security-unverified-token.json267 B
- reviews/simulated-reviewer-b/silent-default-flip.json229 B
- reviews/simulated-reviewer-b/silent-error-swallow.json258 B
fixtures/
- diffs/concurrency-idempotency-race/change.diff536 B
- diffs/concurrency-idempotency-race/ground-truth.json378 B
- diffs/concurrency-lost-update/change.diff451 B
- diffs/concurrency-lost-update/ground-truth.json375 B
- diffs/edge-empty-average/change.diff356 B
- diffs/edge-empty-average/ground-truth.json345 B
- diffs/edge-missing-header/change.diff324 B
- diffs/edge-missing-header/ground-truth.json340 B
- diffs/logic-inverted-entitlement/change.diff380 B
- diffs/logic-inverted-entitlement/ground-truth.json340 B
- diffs/logic-pagination-boundary/change.diff419 B
- diffs/logic-pagination-boundary/ground-truth.json421 B
- ATTRIBUTION.md947 B
- CHANGELOG.md790 B
27 more files not listed here. See all 67 in the repository.