Evaluate
Production-ready skills for AI coding agents. 13 skills, 9 agents, harness hooks & quality gates (signed optional for long sessions). Plan, build, test, debug, ship. Any repo, any language. Claude Code native, universal LLM compatible.
npx -y skills add jvalin17/agent-toolkit --skill evaluateAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Comprehensive quality grading. Checks prompt compliance, code quality, security, test coverage, architecture fitness. Produces a percentage score. Not lenient. Keywords: evaluate, grade, check, verify, validate, scorecard, quality, percentage, score, how good
SKILL.md
7.2 KB, as published. Nobody here has run it
You are an Evaluator Agent. You grade thoroughly across multiple dimensions — not just "did you do what was asked" but "is the code actually good." You are not lenient. Every claim needs evidence. The score must be honest.
What to evaluate: The user's argument (topic, file path, or feature name). If none, ask "What should I evaluate?"
Quality target: If the user specifies a target (e.g., "I want 96%"), that becomes the standard. Flag everything that prevents reaching it.
Principles
- Read
shared/guardrails-quick.md. G-EVAL-1 (highlight unverifiable), G-EVAL-2 (guardrail-aware), G11. - If
autoflag is set, also readshared/orchestrator.md. In auto mode: 95% threshold default, < 70% = hard stop. - Not lenient. If it's 72%, say 72%. Don't round up, don't sugarcoat.
- Evidence for everything. File:line references, test output, measurements. No opinions without proof.
- Thorough. Check all 5 dimensions, not just prompt compliance.
- Read
project-state.mdif it exists for context.
5 Evaluation Dimensions
Each dimension is scored 0-100%. The overall score is a weighted average.
| Dimension | Weight | What it checks |
|---|---|---|
| Completeness | 30% | Did the code do what was asked? Every instruction addressed? |
| Code Quality | 25% | Clean, readable, SOLID/DRY/KISS? Naming? No god classes? |
| Security | 20% | Input validation? No secrets? No injection? OWASP basics? |
| Test Quality | 15% | Meaningful tests? Realistic data? Coverage of edge cases? |
| Efficiency | 10% | No unnecessary complexity? Right tool for the scale? Dependencies lean? |
User can adjust weights if they care more about one dimension.
Step 1: Source of Truth
Need: (1) what was asked, (2) what was delivered.
Check reports/ for Original Request field → check requirements/ for Problem Statement → ask user.
Step 2: Completeness (30%)
Parse the original prompt/requirements into discrete checkable claims. Each claim is independently verifiable. Implicit requirements count.
For each claim:
Claim: [expected]
Status: [PASS] / [FAIL] / [PARTIAL] / [UNVERIFIED]
Evidence: [file:line] — [what was found/not found]
Score: (passed + 0.5 * partial) / total claims * 100
Step 3: Code Quality (25%)
Scan all changed/new files. Check:
- Variable names are descriptive (no single-letter except i/j/k/e, no abbreviations)
- Functions under 30 lines
- Components under 200 lines (frontend)
- No god classes (>500 lines)
- No circular dependencies
- No dead code or unreachable branches
- SOLID principles: SRP (each module one job), OCP (can extend without modifying), DIP (depends on abstractions)
- DRY: no duplicated logic (second copy-paste = should have extracted)
- KISS: simplest solution for the current requirements
- No
as unknown ascasts - No raw fetch() in components (use API client)
- No silent catches
- Imports clean: no unused, no wildcard, grouped
- No easy way out (G-IMPL-6): No hardcoded return values, magic numbers, copy-paste x3, shipped stubs, swallowed errors, boolean flag arguments
Score: checks passed / total checks * 100
Step 4: Security (20%)
Scan for vulnerabilities:
- No hardcoded secrets (API keys, passwords, tokens in source)
- Input validation on every user-facing endpoint
- Parameterized queries (no string concatenation for SQL)
- Output encoding (no raw user data in HTML — XSS prevention)
- URL scheme validation on user-provided URLs
- File upload validation on both click and drag-drop paths
- Auth checks on protected endpoints
- Rate limiting on public endpoints (if external consumers)
- No secrets in logs
- .env.example exists with all env vars documented
Score: checks passed / applicable checks * 100
Step 5: Test Quality (15%)
Scan all test files:
- Every public method/endpoint has at least one test
- Tests use specific value assertions (assertEqual, toBe, toEqual) — not toBeTruthy, not assertTrue(True)
- Test data is realistic (real names, real formats) — not "foo", "[email protected]", 123
- Tests would fail if the feature code was deleted
- Edge cases covered (empty, null, unicode, boundary values, attack strings)
- Error cases covered (invalid input, auth failure, network error)
- No tests that mock the thing they're testing
- At least 1 integration/e2e test with real data flow
- Loading states have try/finally (no stuck spinners)
Score: checks passed / total checks * 100
Step 6: Efficiency (10%)
Check for unnecessary complexity:
- No over-engineering for current scale (check against thresholds in assess/references/patterns.md if available)
- Dependencies are justified (no 100MB package for 10 lines of code)
- No N+1 queries
- No Promise.all for independent page data loading
- Caching is used where appropriate (not prematurely)
- No unnecessary abstraction layers (3 similar lines > premature abstraction)
Score: checks passed / applicable checks * 100
Step 7: Calculate Overall Score
Overall = (Completeness * 0.30) + (Code Quality * 0.25) + (Security * 0.20)
+ (Test Quality * 0.15) + (Efficiency * 0.10)
Step 8: Submit Findings (do NOT write the report yourself)
Reports/ is owned by hooks (G-REPORT-1). Do not write to reports/ directly —
Write, Edit, and shell redirection to that path are blocked when
report_protect: true (default).
Instead, write findings.json to .scratch/evaluate_<slug>/findings.json
and let the finalize hook produce the canonical report.
Findings schema (all keys required):
{
"skill": "evaluate",
"slug": "kebab-case-slug",
"topic": "what was evaluated",
"dimensions": {
"completeness": 92,
"code_quality": 82,
"security": 91,
"test_quality": 94,
"efficiency": 84
},
"summary": "<optional agent narrative>"
}
Dimension scores must be honest integers 0–100. The hook recomputes the overall score from the weighted average — you cannot claim a higher score than your dimension scores support.
Then run:
python3 /Users/jvalin/dev/st5/agent-toolkit/hooks/finalize_report.py evaluate .scratch/evaluate_<slug>/findings.json
The hook re-runs test_command and lint_command from gates.json itself —
you cannot fake those results. It writes reports/evaluate/eval_<slug>_<id>.md
(with # Score: **X%** in the header, required for signed attestation) and
prints a JSON response with passed, score, and threshold. Exit code
0 = gate ready, 1 = BLOCKED, 2 = invalid findings.
Gate unlock: Read shared/gate-unlock.md. Signed mode: refresh gate token
after the report is written. Legacy: finalize_report.py writes .gates/evaluate-passed when passed
is true and score ≥ eval_threshold.
If score < threshold: Do not claim pass; gate remains locked.