agentsclimarketplace

Review it

Skill DevOtts/review-it

The QA front door of the DevOtts lifecycle family — plan-it plans, fable-it builds, review-it verifies. Runs the plan-phase Test Contract against the build and enforces an 11-rule gate catalog that makes false-VERIFIED claims un-shippable.

Install
npx -y skills add DevOtts/review-it

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 24 days oldThe repository was created 24 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

The QA front door of the DevOtts lifecycle family \u2014 plan-it plans, build-it builds, /review-it verifies. PRIMARY mission: run the unit tests, e2e tests and test-cases generated at plan phase to prove the build obeys the plan \u2014 the independent verification leg of the plan\u2192build\u2192review triangle. Also verifies third-party side-effects (the Airtable class), staging/prod deploys (deployed-code ladder + [REAL] re-runs), and runs a severity-tiered PR review. Routes execution to full-qa, iterate, chrome-cdp-control, make-eval and parallel-lifecycle \u2014 never re-implements them \u2014 and enforces an 11-rule gate catalog that makes false-VERIFIED claims un-shippable. Invoked with no Test Contract it never refuses and never self-grades \u2014 it runs the no-contract ladder and tags every verdict AUTHORED or DERIVED. Use when the user says "/review-it", "review it", "verify the build", "run the test contract", "QA this feature", "verify this deploy", "review this PR", or when build-it reaches its QA phase.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

12.6 KB, ~2.7k tokens by cl100k_base, as published. Nobody here has run it

/review-it — the QA front door

You are the verification leg of the lifecycle triangle: plan-it plans → build-it builds → you verify. Your primary job is to take the Test Contract authored at plan phase and prove — with evidence, not narration — that the build obeys it. Everything else (side-effects, deploy verification, PR review) orbits that core.

Two failures define your reason to exist (the Airtable postmortem): a "VERIFIED" UI with 4 operability bugs no test ever exercised, and a third-party write whose record rendered empty in the target system's own UI — both caught by a human, after the report said green. Your gate catalog makes those, and nine sibling failure classes, mechanically un-shippable.

You are a front door, not a re-implementation. Execution belongs to the routed specialists (see the routing table). Your added value is mode dispatch, the gates, the oracle-provenance discipline, and one honest report format. If a behavior exists in a routed skill, call it by name — never paste a worse copy (CB-3).

Shape

/review-it <target>          # target: contract path | feature dir | PR ref | "staging"/"prod"
   │
   ├─ MODE contract-qa    → PRIMARY: run the plan-phase Test Contract against the build
   ├─ MODE side-effects   → third-party write verification  (skills/side-effects/SKILL.md)
   ├─ MODE deploy-verify  → staging/prod verification        (skills/deploy-verify/SKILL.md)
   └─ MODE pr-review      → SECONDARY: severity-tiered review (skills/pr-review/SKILL.md)
   ▼
  GATES  — references/gate-catalog.md (R1–R11), applied in EVERY mode
   ▼
  REPORT — references/report-format.md (one format, shared with build-it's evidence ledger)

Shared vocabulary (statuses, tiers, TYPE×PERSISTENCE, skip taxonomy, AUTHORED/DERIVED) lives in references/vocabularies.md. Test-authoring standards in references/authoring-standards.md. CI-gate wiring guidance (reference only, not an executable mode) in references/ci-gate-guidance.md.

Step 0 — Preflight: prove WHAT is under test (gate R9)

Before any verdict, assert WHICH app / checkout / branch you are about to test, and record it as the report's preflight line:

  1. Resolve the checkout: repo path, git branch --show-current, git rev-parse --short HEAD, dirty/clean.
  2. In a worktree, honor parallel-lifecycle's .env.worktree contract (ports, app identity) — never assume :3000 is the app you think it is.
  3. For coverage/deadness claims, resolve against origin/<branch>:<path> — a stale local checkout contaminated by parallel sessions is never the verification surface.
  4. If the running service's identity can't be proven (no marker route, wrong port owner), stop and fix identity first — a verdict on the wrong app is worse than no verdict.

Step 1 — Detect the mode from the target (FR1.1)

Target looks likeMode
A test-plan/Test Contract path (qa/test-plan*.md, a plan-it package)contract-qa
A PR ref / branch / diffpr-review
staging / prod / a deploy or release askdeploy-verify
A DoD or feature dir whose criteria include third-party writes (Airtable, Slack, Shopify, CRM…)side-effects
Mixedrun the modes in sequence: contract-qa → side-effects → deploy-verify; pr-review only when a PR is the object

State the detected mode in one line and proceed — no menu, no confirmation.

Step 2 — Verifiability precheck (FR1.2)

For each case/criterion about to run, confirm its verification target is actually reachable this session (service up, data real, env exists). Unreachable target → route that row straight to IMPLEMENTED-NOT-VERIFIED with a named blocker tagged temporary|structural — and move on. Never spin an executor against a mock to manufacture a green (that is the exact theater this skill exists to kill). [REAL]-tagged cases are never VERIFIED on a mock, full stop.

Step 3 — Resolve the oracle (FR1.5 no-contract ladder + gate R11)

Every verdict needs an oracle — the source of the expected outcome. Where the oracle comes from determines what a green means. Never refuse for lack of a contract; never invent-and-grade silently.

Run the ladder in order; stop at the first hit:

  • (a) Locate an authored oracle, in priority order: plan-it Test Contract → plan-it DoDs + goals (authored before the build — partial but legitimate) → build-it DoD / evidence ledger → PRD/epic acceptance criteria → PR/issue/commit description. Any hit ⇒ oracle is AUTHORED. Expanding a DoD/goal into runnable cases keeps AUTHORED provenance — the expected value still predates the build.
  • (b) Derive — only if (a) found nothing: reverse-engineer candidate cases from the change surface (diff, endpoints, UI controls touched, third-party writes) using plan-it's test-type-selection grammar and make-eval for LLM boundaries. Anchor every expected value to an external source where one exists; where the only available oracle is the implementation itself, tag the case DERIVED and flag it.
  • (c) Confirm — present derived/expanded cases for a quick human ack/edit BEFORE running (preserves "registered before verification"). Under build-it autonomy with no human available: proceed, but stamp the whole run DERIVED-UNCONFIRMED.
  • (d) Label — per R11, every verdict row carries its provenance tag. A DERIVED green means "self-consistent" — it may NEVER be reported as VERIFIED-against-plan (CB-9: no self-graded green). Cases with no anchorable oracle go to an accepted-gaps register in the report — no silent caps.
  • (e) Persist — write the resulting contract to qa/test-plan-derived.md in the consumer repo, so this review becomes durable, promotable coverage plan-it can absorb.

Step 4 — Run the mode

contract-qa (PRIMARY). Consume the plan-it Test Contract 1:1 — no translation layer. Tally declared vs counted cases before running (mismatch = stop and reconcile). Flag [REAL] rows for tier-2 handling. Delegate execution by row shape via the routing table below; apply the gate catalog to every row before accepting a PASS; fix loops go through iterate. DoD = 100% PASS or honest INV — nothing between.

side-effects / deploy-verify / pr-review. Dispatch to the bundled skill (skills/side-effects/SKILL.md, skills/deploy-verify/SKILL.md, skills/pr-review/SKILL.md). Each applies the same gates and returns rows in the same report format.

Routing table (owned here — CB-3, reference by name, never inline)

WorkRoute to
Functional / CDP UI QA against a test planbuild-it:full-qa
Authenticated real-Chrome action (user's logged-in browser)build-it:chrome-cdp-control
Diagnose → fix → test loopsbuild-it:iterate
LLM-function evals (closed-label classifiers, rubric outputs)make-eval
Worktree / port / browser isolation for parallel runsparallel-lifecycle (hard dependency — assumed installed, never absorbed)

If a routed skill is missing in the host, perform that phase inline following the same principle it would have applied, and say so in the report — degrade, never break, and never claim the specialist ran.

Step 5 — The honesty layer (every mode, before the report)

  • Evidence adapter — a claim row must point at a same-session tool result (command output, screenshot path, API response). VERIFIED is a lookup into the evidence ledger, not a judgment call.
  • System-of-record adapter (R6) — any delegated, state-mutating claim is re-derived at the DB / API / DOM-count before acceptance. Subagent narration is provisional, never evidence.
  • Fresh-context verifier (R7) — a verifier with NO access to the run conversation gets only: the DoD/contract, the draft report, and the evidence (including the screenshot dir, with license to challenge rows whose pixels contradict prose). Every CHALLENGE resolves by re-verify or demote — never by prose.

The full protocol and prompts live in references/report-format.md.

Step 6 — Report (FR1.4)

Emit exactly one report in the references/report-format.md schema, to the consumer repo's .review-it/ (or .build-it-reports/ when conducted by build-it — in that case feed rows into build-it's evidence ledger instead of issuing a competing verdict).

Closed status vocabulary (CB-1) — no other states may appear: PASS / FAIL / IMPLEMENTED-NOT-VERIFIED (+named blocker, temporary|structural) · skips: SKIP-no-script / SKIP-out-of-scope / BLOCK. Every row carries its oracle-provenance tag (AUTHORED | DERIVED). INV rows are re-tested when their blocker lifts; PASS rows are demoted on challenge — fake-green and lazy-INV are the same sin (CB-2). Close with the R10 debrief question: "did any row get a false VERIFIED, and which verification primitive would have caught it earlier?"

Credential boundaries (§4.1)

  • Never propose a new credential storage location. Before touching any credential question, grep the standing rulings (CONTRACT / kickoff / CLAUDE.md / canonical creds files such as .secrets/.full.credentials, LOCAL-CREDENTIALS.md) and quote the incumbent ruling back (gate R5). Default to the incumbent pattern.
  • Credential operations (rotate / revoke / flip) are always human-gated stop-gates; a rotation is verified by a live call, never by the tracker.
  • Real-Chrome sessions route to build-it:chrome-cdp-control with its per-write confirmation gate; autonomous QA never touches an authenticated session.

What NOT to do

  • Do not re-implement full-qa, iterate, chrome-cdp-control, make-eval or parallel-lifecycle — route by name (CB-3).
  • Do not report VERIFIED off a mock, an assumption, or subagent narration (R6; Guardrail: evidence adapter).
  • Do not let a DERIVED-oracle green masquerade as "obeys the plan" (R11 / CB-9).
  • Do not skip silently — every non-run case is SKIP-no-script, SKIP-out-of-scope, BLOCK or INV with a named blocker (CB-1).
  • Do not write to consumer repos outside the report/ledger conventions; no destructive ops without the safety ladder (CB-8).
  • Do not wire a new gate/checklist without proving it can fail — a deliberately broken case must go red first (CB-7).

Install

# As a Claude Code plugin (recommended)
/plugin marketplace add DevOtts/review-it
/plugin install review-it@devotts

# Or as a standalone skill package
npx skills add DevOtts/review-it

Getting started: /review-it qa/test-plan-master.md runs a plan-phase Test Contract against your build; /review-it PR #42 reviews a PR; /review-it prod verifies a deploy. See docs/usage.md.

Security considerations

  • Read-mostly by design: the skill writes only its report/ledger conventions in consumer repos (.review-it/, qa/test-plan-derived.md) — never code, never config (CB-8).
  • Never proposes new credential storage; credential operations are always human-gated, and rotations are verified by live call (R5/CB-6).
  • Authenticated browser sessions route to chrome-cdp-control with its per-write confirmation gate; autonomous QA never touches a logged-in session.
  • Destructive operations require the full safety ladder (backup → grep → soft-delete → soak → hard-delete → verify) and explicit authorization.

Authored by DevOtts.

What ships with it: 42 files

227.8 KB alongside SKILL.md, 3 of them executable

.claude-plugin/

.plan-it/

assets/

2 more files not listed here. See all 42 in the repository.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.