Challenge plans
Before you execute a drafted plan/spec/design doc, run a multi-agent adversarial cross-review to surface the flaws that cause downstream rework, aggregating "evidenced, cross-family-verified" objections into a verdict. Use when the user asks to "review this plan/spec", "can this approach be executed", "poke holes / adversarial review / QA this", "harden before executing", or when an agent is about to hand a drafted decision/QA back to the user — first run this skill and present the cross-review recommendation plus surviving objections. Runs on local subscription CLIs (claude/codex), no per-token API cost. Not for "help me pick among options" — that's the weigh-options deliberation skill.From its SKILL.md
npx -y skills add hiadrianchen/challenge-plansAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
3 things to look at
- reads credentialsReads from 1 credential source: `CP_BYO_<n>_TOKEN`.
- 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 7 commands, including `challenge-plans doctor` and 6 more.
SKILL.md
7.6 KB, ~1.7k tokens by cl100k_base, as published. Nobody here has run it
challenge-plans
Harden a plan/spec in multi-agent adversarial review before execution, to reduce rework. Slots into writing-plans → challenge-plans → executing-plans.
When to use (routing signals)
- Input is a single drafted artifact + intent "review / find flaws / can this execute / harden / QA" → use this skill. Pick
--typeby what the artifact is:- a spec / design doc / PRD you're about to build →
--type spec - any plan with steps to execute (dev or not — a trip, a launch, a hire) →
--type plan - a code change (
git diff) →--type diff - a decision already made (an ADR / "we chose X because Y" — a tech-stack pick, a vendor, a hire) →
--type decision. Audits the choice itself: skipped alternatives, weak evidence, sunk-cost reasoning, irreversibility.
- a spec / design doc / PRD you're about to build →
- An agent has finished something and is about to ask the user to decide or QA → run this first and present the cross-review recommendation.
- Input is ≥2 options still open to choose among → use the sibling
weigh-options(deliberation/voting), not this. (--type decisionis the opposite: one option already chosen, audited after the fact.)
Run
# Installed from PyPI (pip install challenge-plans) — the console command is available directly.
# Already installed but on an old version? update: pip install -U challenge-plans (pipx: pipx upgrade; uvx: append @latest).
challenge-plans doctor # check adapter login state first
challenge-plans run <artifact> --type spec --profile standard --sink markdown # review a plan/spec
challenge-plans run <artifact> --type spec --profile standard --sink markdown --lang zh # localized output
# code-diff gate: git diff > change.diff && challenge-plans run change.diff --type diff --sink markdown
# From a source checkout instead (not pip-installed): PYTHONPATH=src python3 -m challenge_plans.cli <args>
--type spec|diff|plan|decision.diffreviews a rawgit diff;planreviews ANY plan (a trip, a launch, a hire — not just dev specs) with domain-neutral failure types (missing success criteria / ignored constraint / unaddressed risk / sequencing gap / unstated assumption / goal misalignment / irreversibility / no fallback) and feasibility·risk·goal-alignment lenses;decisionaudits a choice already made (ignored alternative / weak evidence / unstated assumption / sunk-cost bias / unaddressed downside / irreversibility / no review trigger / misframed problem) with alternatives·evidence·reversibility-cost lenses. All run the same verdict pipeline.--profile fast|standard|deep,--sink stdout|markdown,--enforce(request_changes/inconclusive/schema_invalidexit non-zero;discuss/approveexit 0);--strict(hard gate — only a cleanapprovepasses); default advisory exits 0.--lang <code>(defaulten): write the review prose in the user's language, e.g.--lang zh. Set this from the user's language so the whole review comes back localized; JSON keys / enums /L12-15anchors stay stable. Equivalent to exportingCHALLENGE_PLANS_LANG.- Output: a 6-state verdict + surviving objections.
[sev✓]= cross-family Verifier-confirmed, may hard-gate;[sev?]= unverified, advisory only.
If no backend is ready
challenge-plans needs at least one logged-in subscription CLI (it has no model of its own and uses no API keys). If doctor shows nothing ready, don't retry blindly — ask the user, then route:
- Has a Claude or ChatGPT subscription, but the CLI is missing / logged out → walk them through the exact step
doctorprints (install it, orclaude→/login, or sign in tocodex). - No subscription yet but wants one → point them to subscribe (Claude Pro/Max or ChatGPT), then install + log in the CLI.
- No subscription and doesn't want one → explain challenge-plans cannot run without one, and stop — don't loop.
doctor already prints the per-backend fix plus this guidance; surface it to the user rather than failing silently.
BYO backends (optional). The user can register extra Anthropic-compatible endpoints (GLM/Kimi/a proxy) via env: CP_BYO_<n>_BASE_URL + CP_BYO_<n>_FAMILY + CP_BYO_<n>_TOKEN (all required; _MODEL optional; _1/_2/… for several). Two caveats to relay honestly: (1) the declared family is user-declared, not verified — confirmations through it render ✓(user-declared family) and diversity built on it is flagged, so never present a BYO pairing as the verified builtin claude+gpt cross-family guarantee; (2) the token goes only into that backend's subprocess and never appears in any output — but advise the user to set it via a secrets manager / leading-space export rather than plain shell history. A backend missing one of the three vars is skipped with a warning (never a fallback to subscription auth).
If a backend is too old / a run degrades or errors opaquely
A backend CLI that is logged in but out of date is a distinct failure from "logged out": login status still passes, but a real call is rejected server-side (observed: codex 400 "the model requires a newer version of Codex"), so a voter reports exit_nonzero and the run silently drops to a single family. A run can't cheaply tell this apart mid-flight — but doctor now can: it sends a real minimal call per backend, so a too-old codex reads unsupported_version → update Codex CLI: npm i -g @openai/codex@latest instead of a false ready.
So when a run errors, comes back single-family unexpectedly, or a voter shows exit_nonzero: run doctor and update any unsupported_version backend before retrying — don't just report the error to the user. Updating the backend CLI (npm i -g @openai/codex@latest, or update Claude Code) is part of the standard fix path, not a dead end to hand back.
Presenting to the user
Surface the verdict + surviving objections (✓ verified vs ? unverified) + missing required fields as "my cross-review recommendation", then let the user decide — rather than handing them a bare decision. See README.md for the full picture.
Composing with planning skills
- superpowers (
writing-plans → executing-plans): afterwriting-planssaves a plan file (defaultdocs/superpowers/plans/<date>-<feature>.md— read the actual path), runchallenge-plans run <plan> --type specbeforeexecuting-plans. It occupies the same pre-execution review seam as superpowers' built-inplan-document-reviewer, but as a multi-CLI cross-family pass. Route surviving objections back into the plan, then execute. - grill-me (mattpocock/skills): complementary and earlier — it interactively aligns the user while the plan forms (no file output). Run challenge-plans after a written plan/PRD exists.
- Nothing auto-invokes challenge-plans; the calling agent wires it into the seam and chooses
--typefrom the routing signals above.
What ships with it: 25 files
293.7 KB alongside SKILL.md, 12 of them executable
docs/
examples/
- decision-sample.md1.4 KB
- options.yaml471 B
- plan-sample.md958 B
- spec-sample.md582 B
src/
- challenge_plans/adapters.pyruns28.3 KB
- challenge_plans/cli.pyruns27.0 KB
- challenge_plans/config.pyruns4.7 KB
- challenge_plans/deliberation.pyruns9.6 KB
- challenge_plans/engine.pyruns19.5 KB
- challenge_plans/__init__.pyruns705 B
- challenge_plans/preflight.pyruns4.6 KB
- challenge_plans/prompts.pyruns4.5 KB
- challenge_plans/rubric.pyruns7.2 KB
- challenge_plans/schema.pyruns11.3 KB
- challenge_plans/verifier.pyruns7.0 KB
tests/
- test_challenge_plans.pyruns94.4 KB
- CONTRIBUTING.md1.3 KB
- .gitignore86 B
- LICENSE11.1 KB
- pyproject.toml1.1 KB
- README.md16.8 KB
- README-zh.md15.1 KB