agentsclimarketplace

Eval review

Skill easyinplay/harnessed/workflows/verify/eval-review

AI coding harness composition orchestrator — manifest-described upstreams, composition skill workflows. Apache-2.0.

Install
npx -y skills add easyinplay/harnessed --skill eval-review

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Stage ④ verify sub-workflow — GSD /gsd-eval-review AI phase eval 覆盖审计 (has_ai_phase 触发, 可选 conditional; pairs with plan 侧 gsd-ai-integration-phase AI-SPEC eval strategy). schema_version: harnessed.workflow.v3 with disciplines_applied (6 default) + tools_available (gsd-eval-review) + 1 phase (gate ref has_ai_phase conditional)。 Triggered by slash command `/verify-eval-review` after `harnessed setup`.

SKILL.md

5.2 KB, as published. Nobody here has run it

verify-eval-review workflow (v3)

Overview

1-phase sub-workflow auditing AI phase eval coverage (v13.0 P42 upstream re-sync — D-04 Stage ④ Verify conditional sub + GSD eval-review wire)。Pairs with the plan-side gsd-ai-integration-phase (AI-SPEC.md eval strategy): verify 侧回查实现是否真覆盖规划的 eval 维度,产出 EVAL-REVIEW.md (逐维度 COVERED/PARTIAL/MISSING)。

phaseidupstreammodelcapabilitygate
101-eval-reviewgsdsonnet{{ capabilities.gsd-eval-review.cmd }}judgments.stage-routing.verify-eval-review-aiphase.fires

Per-phase config loads from workflows/verify/eval-review/workflow.yaml; engine 4-level gate resolver evaluates phase.has_ai_phase == true via expr-eval — true 则 invoke GSD /gsd-eval-review (eval 覆盖审计 → EVAL-REVIEW.md), false 则 skip。

Capability refs

Sister workflows/capabilities.yaml entries:

  • gsd-eval-review — Bucket 2 special-purpose (impl: gsd, cmd: /gsd-eval-review, fires_when: has_ai_phase)

Gate ref

Sister workflows/judgments/stage-routing.yaml:

  • verify-eval-review-aiphase.firesphase.stage == 'verify' and phase.has_ai_phase == true

How to invoke

!harnessed checkpoint intent verify-eval-review

The banner above (when present) means this invocation is REGISTERED with the engine (an intent marker) — not yet compliant: the steps below (prompt → spawn → checkpoint complete) resolve it, and a per-turn <workflow-intent> reminder persists until they run.

The numbered sequence below is the state machine — execute it with Bash. Do NOT improvise an equivalent flow from the Overview above: freelancing bypasses the engine (no ledger, no evidence guard). harnessed gives you the spawn-ready prompt; YOU spawn the subagent with a CC-native Task / Agent tool (keeps the session responsive + lets clarification round-trips reach the user).

Do NOT pipe to harnessed run verify-eval-review — that is the CI/headless path (in-process SDK spawn that blocks the session inside Claude Code).

  1. Bash: harnessed prompt verify-eval-review --task "$ARGUMENTS" --json → parse {prompt, max_iterations, model}.
  2. Spawn a CC-native subagent (Task / Agent tool) with that prompt and model, then drive delivery with harnessed's own completion gate:
    • on return, write the subagent's final output to a file and run harnessed checkpoint complete verify-eval-review --result-file <path> — it is fail-closed on the declared artifacts, the TDD boundary, and the verbatim <promise>COMPLETE</promise>.
    • if it blocks, run harnessed checkpoint fail verify-eval-review --failing-tests <n> to record the attempt; it prints BUDGET-EXHAUSTED / NO-PROGRESS / BREAK-LOOP when a stop condition is reached.
    • respawn ONLY while none of those three has fired. Any one of them means stop: re-scope the subtask, fix the blocker, or escalate to the user. Never respawn past a stop directive.
  3. If the output contains STATUS: NEEDS_CLARIFICATION + a question list: STOP, relay them verbatim via AskUserQuestion, append the answers to the spec, then re-spawn the same sub.
  4. On <promise>COMPLETE</promise>: write the subagent’s final output to a file, then Bash harnessed checkpoint complete verify-eval-review --result-file <path> --summary "<one-line>". Fail-CLOSED — it blocks unless every declared artifacts_expected file exists, the TDD boundary passes (non-empty evidence / both the red and green sides present / the test file was not deleted), and the result carries a verbatim <promise>COMPLETE</promise> (or a structured COMPLETE status). --result <text> is the inline variant; --result-file wins and is quoting-safe on Windows. --force records an audited override (evidence_status=overridden) — it does not silently pass.
  5. If the complete gate blocked: Bash harnessed checkpoint fail verify-eval-review --failing-tests <n> to record the attempt. It prints BUDGET-EXHAUSTED / NO-PROGRESS / BREAK-LOOP once a stop condition is reached. Respawn ONLY while none of those three has fired; any one of them means STOP — re-scope the subtask, fix the blocker, or escalate to the user.
<!-- harnessed-generated:v4.12.0 -->

References

  • D-04 Stage ④ Verify conditional sub 分解
  • v13.0 P42 upstream re-sync — GSD eval-review wire (pairs with plan 侧 gsd-ai-integration-phase)
  • workflows/capabilities.yaml — gsd-eval-review
  • workflows/judgments/stage-routing.yaml — verify-eval-review-aiphase trigger
  • workflows/verify/qa/workflow.yaml — sister conditional-sub pattern (has_ui_changes gate)

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.