agentsclimarketplace

Concilium

Skill raichominev/concilium

Cross-model adversarial review skill for Claude Code: Fable orchestrates, GPT (sol/terra/5.5 via codex CLI) probes, verdicts get ratified. ChatGPT subscription auth, no API key.

Install
npx -y skills add raichominev/concilium

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 21 days oldThe repository was created 21 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Adversarial cross-model review for hard, load-bearing tasks — combining frontier models: the Claude session (Opus 5 or Fable 5 as the intended orchestrator) hands a claim, diff, or result to an OpenAI model (gpt-5.6-sol / gpt-5.6-terra / gpt-5.5, via the codex CLI on ChatGPT-subscription auth, no API key), which probes it with falsification attempts and PROPOSES a verdict; the orchestrator checks the probe and RATIFIES. Use whenever the user wants a second opinion from a different model, a cross-model or concilium review, adversarial verification of a research claim, benchmark number, or diff, says "have GPT/codex check this", wants codex set up as a reviewer, needs to switch codex models mid-session (park-and-resume), is tiering work across codex models, or wants to LOOP/iterate review rounds until a disputed claim converges.

SKILL.md

13.0 KB, ~3.1k tokens by cl100k_base, as published. Nobody here has run it

Concilium — cross-model adversarial review

A second, different model reviews your (or the user's) claims adversarially. Different model lineage means different blind spots — that's the value. The reviewer PROPOSES; the calling session RATIFIES. Never let either side's confidence substitute for evidence.

Designed to be orchestrated from Claude Code — Opus 5 and Fable 5 are both first-class ratification seats (measured at chair parity on a blind outcome-prediction benchmark; any Claude model can drive the loop, but the ratifier should be one of the two). The GPT side (sol/terra/5.5 via codex) does the independent probing and mechanical execution — and that cross-family seat is load-bearing: it measurably catches what same-family chairs jointly miss.

Prerequisites (check once per environment)

  1. codex login status → must say "Logged in using ChatGPT" (subscription OAuth — an API key is NOT needed and a subscription can NOT be used as one; don't attempt proxy/router bridges).
  2. Discover available models: codex debug models or ~/.codex/models_cache.json. If a model errors "requires a newer version of Codex", run codex update and retry.
  3. First time in a new environment, run the calibration bootstrap (references/setup.md) before trusting verdicts: a known-truth reasoning test, then one simple real task, then (optionally) a head-to-head to pick tier models.

Tier matrix (defaults are current-day models — override per installation)

TierDefaultEffortUse for
Researchflagship (e.g. gpt-5.6-sol)highopen review rounds, adversarial verification
Mechanicalprev flagship (e.g. gpt-5.5)mediumverify a known claim with one probe
Runnercheap tier (e.g. gpt-5.6-terra)lowexecute-and-report: run a script, babysit an import

Runner tasks are NOT reviews — skip the wrapper and call codex directly: codex exec -m <cheap-model> -c model_reasoning_effort=low [-s read-only unless it writes] "<task>" < /dev/null Always close stdin on direct non-interactive calls (bash < /dev/null; PowerShell $null | codex …) — an open non-TTY stdin blocks codex forever on "Reading additional input from stdin...", and the orphaned process survives the parent shell's timeout (pitfall #10; the wrappers are immune — they pipe via stdin).

Running a review

Use the bundled wrappers. They load the shared review contract from references/contract.md (single source of truth — falsification probe, alternative explanation, caveat, verdict-proposal, schema/encoding rules; edit the contract there, never in the scripts) and add provenance stamping. Both wrappers are functionally identical; pick by platform:

Linux / macOS (bash):

  • Claim: scripts/concilium-review.sh claim "<claim>"
  • Diff: scripts/concilium-review.sh diff [base-branch]
  • Config via env: MODEL, EFFORT, MECHANICAL=1 (mechanical tier), REPO_DIR, PROJECT_RULES (rules file path).
  • First use after clone: chmod +x scripts/concilium-review.sh.

Windows (PowerShell 5.1+):

  • Claim: powershell -ExecutionPolicy Bypass -File scripts/concilium-review.ps1 -Claim "<claim>" [-Mechanical] [-RepoDir <path>] [-ProjectRules <file>]
  • Diff: ... -Diff [-Base <branch>] — reviews the working-tree diff of -RepoDir.

Operational rules (each one is a measured failure — the why is in references/pitfalls.md):

  • Run in background with a full ~10 min timeout from the FIRST call. Real reviews take 5–15+ min at high effort; a foreground timeout kills them mid-probe.
  • Prefer a fresh session over resuming a timed-out one. Long resumed chains hit context compaction — the reviewer's early careful reading gets lossy-summarized before the final, consequential step.
  • Never bare-resume. codex exec resume silently resets model AND sandbox to the user's config.toml defaults. If you must resume (or want to switch models mid-session), re-pin everything: codex exec resume -m <model> -c sandbox_mode="read-only" -c model_reasoning_effort=<tier> <session-id> - Flags go BEFORE the positional session id. The key is sandbox_mode-c sandbox=... is silently ignored, and there is no -s flag on resume. Cross-model resume retains context.
  • The reviewer is a full agent, not a chatbot — read-only sandbox blocks file writes, not read commands or DB SELECTs. Everything it reviews goes to the second model's provider.
  • Watch progress live, don't wait blind — and monitor the right stream. The contract (rule 9) makes the reviewer emit STATUS: one-liners as it works, and codex writes progressively — but the streams split (verified live): with 1> out 2> err, the final five blocks land on stdout while the streaming transcript (banner, STATUS lines, tool calls) goes to stderr. Point a tail/monitor at stderr for progress + failure signatures; read stdout for the verdict. Caveat: PowerShell 1>/2> redirects write UTF-16 — decode accordingly (or redirect through a UTF-8-forcing step) before grepping.
  • A blind round needs structural isolation, not an instruction. When the round must be unprimed (a blind eval, a framing-critical blind-first pass per request-template), run the reviewer in a clean directory with auto-rules bridging OFF (-NoAutoRules / NO_AUTO_RULES=1): a model carrying project context and told to "answer from the packet alone" measurably still uses that context (pitfalls #16–17).

Ratification protocol (the calling session's job)

The reviewer returns five blocks: PROBE / ALT / CAVEAT / VERDICT-PROPOSAL / PHASE-LOG. Before relaying or acting:

  1. Read the actual probe (the query/commands), not just the prose summary.
  2. Extremal results are a tripwire: 0% or 100% on a first attempt usually means a wrong join key, wrong scope, or wrong table — not a discovery. Verify the probe's load-bearing step yourself before accepting it.
  3. Scope-check disagreements: two probes can both be factually right at different scopes (one table vs DB-wide, one source vs all sources). Name the scope before comparing numbers.
  4. Distinguish refuted / stale / incomplete. "The numbers differ today" does not mean the claim was wrong when written — check history/timestamps before saying "refuted".
  5. Assign the final verdict tag yourself: [V-code] (verified vs source, cite file:line) / [V-db] (read-only query, cite it) / [V-probe] (re-runnable script) / [C] (unverified) / [X] (refuted — name what supersedes it). The proposal is input, not the answer.
  6. Weigh agreement by lineage. Same-family confirmation (a Claude chair agreeing with a Claude orchestrator) is weak evidence — same-lineage chairs measurably share wrong answers, down to independently producing the identical wrong inference. A cross-family confirmation or refutation outweighs any count of same-lineage votes; never settle a dispute by majority across chairs that share a lineage.

The concilium loop (iterative rounds)

A single review pass is often enough. But when the reviewer's probe has a gap, or you (the orchestrator) disagree with the proposal on defensible grounds, one exchange isn't a concilium — a council deliberates. The loop runs review rounds until the verdict converges or the dispute is proven genuine. This loop is orchestrated by you, the calling Claude session — it is a protocol, not a script (the ratification step is your judgment; nothing can automate it).

Each round:

  1. Run a review (the wrapper) → get the five blocks → ratify per the protocol above.
  2. Decide the round's outcome and act:
OutcomeConditionAction
ConvergedYou verified the probe's load-bearing step and it holdsSTOP — emit the final tag.
DisputeThe probe has a gap, wrong scope, or you have a specific, evidence-backed objectionWrite this round's PROBE + your objection to a rounds file; run the next round with -PriorRounds/PRIOR_ROUNDS pointing at it.
DryA round adds no new checkable evidence — the reviewer re-asserts, or says (in CAVEAT) it has no new pathSTOP — escalate to the owner as [C]/[POLICY] with the open question. This is the anti-oscillation guard.
CapRound limit reached (default 3) without convergingSTOP — present the state and escalate; a real dispute is a finding, not a failure.

Design rules (they follow directly from the pitfalls):

  • Fresh session per round — never a resume chain. The loop is exactly the "long chain" that pitfall #3 warns about; carry context forward via the -PriorRounds file, not codex exec resume. Each round starts clean and sees only a compact summary of what was already tried.
  • Every round must add a NEW evidence path. The contract (rule 8) enforces this on the reviewer side; you enforce it on yours — an objection is only worth a round if it's backed by evidence or points at a concrete, checkable gap. "I'm not convinced" is not a round.
  • Ratifier stays fixed (you / Fable); the reviewer can drop tiers as the dispute narrows. Round 1 on the research tier; once it's down to a mechanical check, run later rounds -Mechanical. Each round is a real 5–15 min codex call — the cap and the dry-stop are cost controls, not just correctness ones.
  • Keep the rounds file in durable project storage (not a session temp dir), so the whole deliberation is auditable and the final PHASE-LOG can cite it.

Trigger it when the user asks to "loop", "iterate", "keep going until it's resolved", "have them hash it out", or when a first pass comes back disputed and the stakes justify another round.

Project adaptation

The reviewer sees AGENTS.md, not CLAUDE.md — mind the gap

codex auto-loads AGENTS.md (from the working directory upward), the same way Claude Code auto-loads CLAUDE.md. They are different files: a project with only a CLAUDE.md gives the reviewer none of the ground rules Claude has — it reviews half-blind. Three ways to close it, in order of durability:

  1. Best (project-level): make AGENTS.md exist. Mirror your CLAUDE.md into an AGENTS.md (or make AGENTS.md a short pointer to it), and keep them synced. This helps all codex usage, not just this skill, and is codex's own supported convention.
  2. Automatic (built into the wrappers): CLAUDE.md bridging. When no AGENTS.md is present, the wrapper auto-injects the project's CLAUDE.md (root or .claude/CLAUDE.md) into the contract and prints a notice, so the reviewer isn't missing rules. Disable with -NoAutoRules / NO_AUTO_RULES=1 (e.g. a huge, mostly-workflow CLAUDE.md you don't want in every review).
  3. Curated (explicit): -ProjectRules <file>. Point at a short, hand-picked extract of the safety-critical rules — this overrides auto-bridging. Best for large instruction files where only a slice is relevant to review (invariants, "never touch X", schema quirks).

If you keep both files but let them drift, the reviewer sees the AGENTS.md version — sync them.

Other adaptation

  • If the project keeps a claims ledger, the PHASE-LOG block is a ready-to-paste line (Phase N — <reviewer>(<model>) — <date> — <found> [proposed]); append it only via the project's own hygiene rules (typically: owner or main session, append-only). No ledger → drop the block.
  • Storage: keep probe outputs and frozen samples in a durable project location, never in session-scoped temp dirs (they die with the session).

References

  • references/request-template.md — how to construct the REQUEST you hand in (your side, not the reviewer's): confidence-tag facts (never "do not re-derive" over a conclusion), always mount the repo/DB, license rejecting the frame, and run a blind-first pass for framing-critical rounds. Read before writing any non-trivial request.
  • references/contract.md — the review contract the wrappers send (edit it there; both scripts load it at runtime).
  • references/pitfalls.md — known issues and the rules that counter them (read when a rule seems overcautious, or when debugging reviewer misbehavior).
  • references/setup.md — first-time setup, calibration bootstrap, and the head-to-head method for picking tier models.

What ships with it: 10 files

54.5 KB alongside SKILL.md, 2 of them executable

evals/

references/

scripts/

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.