Quality check eval
Skill mattdweigand-sketch/agent-skills/skills/quality-check-eval
Portable agent skill library for Codex, Claude Code, and other AGENTS-aware tools. Covers eval loops, research, writing, and project hygiene.
npx -y skills add mattdweigand-sketch/agent-skills --skill quality-check-evalAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Sort AI output quality checks into four buckets: automate in code, judge with an LLM, keep for human review, or remove. Use when reviews are too manual, evals are vague, a workflow needs clearer quality gates, or the user wants to decide which AI output checks are worth automating. Not for checking whether a system grades encoded judgment against real outcomes; use cyborg-check for that.
SKILL.md
3.7 KB, as published. Nobody here has run it
Quality Check Eval
Use this when an AI workflow has quality checks, review notes, rubric items, or acceptance criteria, and you need to decide how each one should be handled.
The skill sorts each check into one of four buckets: automate in code, judge with an LLM, keep for human review, or remove.
This is useful when reviews are too manual, evals are vague, or a team is unsure which quality checks are worth automating.
Load references/eval-check-routing.md when auditing. That file owns the stable
recommendation taxonomy, scoring fields, deterministic alternatives, and report
template.
Use This When
Use this skill when the user asks to:
- audit
type: judge, manual, rubric, or subjective eval checks - decide whether to build an LLM-judge runner
- review skipped eval coverage
- rank eval checks by frequency, stakes, and automation value
- convert subjective checks into deterministic checks where possible
Contract
Produces: a markdown audit report in the current repo, plus a short chat summary.
Does not produce: runner implementation, schema migration, scheduler changes, or production writes unless the user explicitly asks after the audit.
Assumptions To State
Before acting, state:
- which repo/worktree you are auditing
- where eval checks live
- where eval logs or recent artifacts live
- what report path you will write
If any of these are unclear and not discoverable from local files, ask one concise question.
Audit Workflow
-
Find eval checks. Use
rgand local file inspection. Look for terms such astype: judge,type: manual,success_criteria,rubric,judge,manual,criteria, andeval. -
Catalog eval checks. Build a table with: check id, parent job/file, type, rubric/check text, expected evidence, and current evaluator path if any.
-
Scan execution signal. Read eval logs, run reports, or recent output artifacts. Count how often each check is skipped, manually reviewed, failed, or would have been relevant. Prefer actual logs over speculation.
-
Check deterministic alternatives. Use the deterministic alternatives in
references/eval-check-routing.md. -
Spot-check real artifacts. Inspect 2 to 5 recent representative outputs. Ask whether each check would have caught a real quality issue. If it would not have fired, call that out.
-
Score each check. Use the scoring fields in
references/eval-check-routing.md. -
Recommend one action. For each check, choose exactly one action from
references/eval-check-routing.md. -
Write the report. Save inside the audited repo. Prefer an existing eval/report folder. If none exists, use
research/orreports/. Use the report template inreferences/eval-check-routing.md.
Report Template
Use references/eval-check-routing.md.
Output Rules
- Be direct. Do not automate checks just because they exist.
- Prefer deterministic checks over LLM judges when the signal can be verified mechanically.
- Separate "valuable review question" from "worth automating."
- Include file paths and line references for important claims.
- End with a short summary: top promotions, auto-conversions, drops, and rough implementation effort.
- Provide a
computer://link to the saved report.