Eval verdict write
Skill 0SxD/advisor-trackable-eval-skills-v01/skills/eval_verdict_write
npx -y skills add 0SxD/advisor-trackable-eval-skills-v01 --skill eval_verdict_writeAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
4.1 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
skills/eval_verdict_write/SKILL.md
Status: live Type: advisor-return skill Inputs: a populated manifest plus the artifacts it references Output: a verdict file at
<workdir>/<task>_EVAL_VERDICT_v1.mdRule reference: seedocs/FLAT_BOOLEAN_RULE.md §7(closing-state semantics) and§8(verdict file shape).
Purpose
Read a manifest the executor has populated with EVIDENCE. For each row,
declare PASS, FAIL, or SKIP. Write a verdict file that downstream consumers
can grep for OVERALL=PASS or OVERALL=FAIL.
Input contract
- manifest_path: path to a populated eval manifest. Every row must be
in a final state:
[x]with non-empty EVIDENCE,[SKIP]with non-empty NOTE, or[ ]with non-empty NOTE describing the blocker. - artifacts_root: directory containing the artifacts the manifest's verification fields reference. May be unreadable (advisor sees only the manifest); the skill handles read-only mode.
- mode: one of
verify(re-run each verification target where possible) ortrust(accept EVIDENCE as recorded).
Output contract
A markdown file at <workdir>/<task_name>_EVAL_VERDICT_v1.md with:
- Top header naming the manifest path it judges.
- Mode line declaring
verifyortrust. - Counts:
PASS_COUNT,FAIL_COUNT,SKIP_COUNT,OPEN_COUNT,OVERALL. - One line per
C-NNNrow in the manifest.
The line format:
C-NNN: PASS | FAIL | SKIP | OPEN -- <one-line basis>
OVERALL is PASS only when:
FAIL_COUNT == 0OPEN_COUNT == 0- Every
[x]row has non-empty EVIDENCE - Every
[SKIP]row has non-empty NOTE
Any other state yields OVERALL=FAIL with the failing condition cited at
the top of the file.
Procedure
-
Read the manifest. Parse each row into
(state, id, phase, subject, verification, expected, evidence, note). -
Run the integrity gates first. If
C-FMT-ROW-SHAPEorC-FMT-NO-STACKfail on the manifest as it stands, the verdict isOVERALL=FAILwith reason "manifest violates its own format rules"; the advisor stops here. -
For each row, validate state.
[x]row with empty EVIDENCE -> FAIL with reason "checked but no evidence recorded".[SKIP]row with empty NOTE -> FAIL with reason "skipped without justification".[ ]row with empty NOTE -> OPEN (not FAIL; the row is incomplete but the executor may not have intended to close it).[ ]row with non-empty NOTE describing a blocker -> SKIP with the blocker as basis.
-
In verify mode, re-run each row's verification target.
- File-existence rows: stat the path.
- Grep rows: re-run the grep with the same pattern and scope.
- Line-existence rows: grep the named log for the named pattern.
- Compare result to
<expected>. Match -> PASS. Mismatch -> FAIL with the diff captured.
-
In trust mode, accept EVIDENCE without re-running. Useful when the advisor session lacks access to the artifacts. The verdict file's mode line declares
trust; downstream consumers know this is a read-only verdict. -
Compute counts and OVERALL.
-
Write the verdict file. Use the same convention as the manifest: one line per row, no narrative, no reasoning beyond the one-line basis.
Failure modes
| Failure | Symptom | Mitigation |
|---|---|---|
Manifest has rows in [ ] state | OPEN_COUNT > 0 | Verdict declares OVERALL=FAIL; downstream stops; executor must close or NOTE the open rows. |
| EVIDENCE field is narrative not fact | Verify-mode re-run finds discrepancy | FAIL row cites the discrepancy. |
| Artifacts not readable | verify mode cannot run | Skill auto-falls-back to trust mode and flags it in the mode line. |
| Verdict file would be empty (no manifest rows) | manifest_path is wrong | Skill stops with explicit "manifest had zero C-NNN rows" error. |
Related
skills/eval_criteria_create/SKILL.mdproduced the input.templates/verdict_template.mdis the empty skeleton this skill writes.tools/lint_no_stack.shruns the integrity gates this skill verifies.docs/FLAT_BOOLEAN_RULE.md §7defines the closing-state semantics.docs/FLAT_BOOLEAN_RULE.md §8defines the verdict file shape.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.