Agent evaluation
VexJoy AI Agent with Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop.
npx -y skills add notque/vexjoy-agent --skill agent-evaluationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
Evaluate agents and skills for quality and standards compliance.
SKILL.md
11.2 KB, as published. Nobody here has run it
Agent Evaluation Skill
Evidence-based quality assessment for agents and skills. The deterministic scorer supplies a 90-point structural precheck; qualitative review covers usefulness and behavior without inventing extra points. Every qualitative finding must cite a file path and line number.
Reference Loading Table
| Signal | Load These Files | Why |
|---|---|---|
| evaluating an entire agent/skill collection | batch-evaluation.md | Loads detailed guidance from batch-evaluation.md. |
| diagnosing recurring structural and content issues | common-issues.md | Loads detailed guidance from common-issues.md. |
| writing single-item or collection evaluation reports | report-templates.md | Loads detailed guidance from report-templates.md. |
| interpreting deterministic scores, JSON keys, or grade boundaries | scoring-rubric.md | Exact contract implemented by score-component.py |
Instructions
Phase 1: Identify Evaluation Targets
Goal: Determine what to evaluate and confirm targets exist.
Read the repository CLAUDE.md first to understand current standards before evaluating anything. Only evaluate what was explicitly requested — do not speculatively analyze additional agents or skills.
# List all agents
ls agents/*.md | wc -l
# List all skills
ls -d skills/*/ | wc -l
# Verify specific target
ls agents/{name}.md
ls -la skills/{name}/
Gate: All targets confirmed to exist on disk. Proceed only when gate passes.
Phase 2: Structural Validation
Goal: Check that required components exist and are well-formed.
Score every rubric category — never skip a category even if it "looks fine." Parse each required field explicitly rather than eyeballing YAML. Record PASS/FAIL with the line number for each check.
Run score-component.py to get deterministic structural scores. It checks frontmatter, referenced paths, pattern and error headings, routing registration, reference-directory presence, workflow structure, and internal links. It does not emit line references or judge content depth, Operator Context, tool semantics, or behavioral quality.
# Deterministic structural checks via score-component.py
python3 scripts/score-component.py agents/{name}.md --json
# or for a skill:
python3 scripts/score-component.py skills/{name}/SKILL.md --json
The JSON output includes results[0].checks with status, earned, max, and detail, plus results[0].total, max_total, and grade. Record these exact keys. Do not refer to earned_points or max_points; those are internal Python attributes, not JSON fields.
See references/scoring-rubric.md for the exact eight checks, 90-point maximum, percentage grade boundaries, optional secret penalty, and JSON contract.
Gate: All structural checks scored with evidence. Proceed only when gate passes.
Phase 3: Qualitative Content Analysis
Goal: Assess whether the component carries useful, accurate, proportionate guidance.
Line counts can describe size, but do not award points for length. More prose is not evidence of better behavior.
# Skill total lines (SKILL.md + references)
skill_lines=$(wc -l < skills/{name}/SKILL.md)
ref_lines=$(cat skills/{name}/references/*.md 2>/dev/null | wc -l)
total=$((skill_lines + ref_lines))
# Agent total lines
agent_lines=$(wc -l < agents/{name}.md)
Check for concrete domain knowledge, stale or contradictory claims, unnecessary bulk, and missing instructions needed to execute the advertised task. Keep these findings outside the deterministic score.
Gate: Qualitative findings cite evidence, or explicitly state that none were found.
Phase 4: Code Quality Checks
Goal: Validate that code examples and scripts are functional.
A script existing on disk does not mean it works — run python3 -m py_compile on every .py file. Search for placeholder text in every file, not just files that "look incomplete."
- Script syntax: Run
python3 -m py_compileon all.pyfiles - Placeholder detection: Search for
[TODO],[TBD],[PLACEHOLDER],[INSERT] - Code block tagging: Count untagged (bare
```) vs tagged (```language) blocks
# Python syntax check
# Syntax-check any .py scripts found in the skill's scripts/ directory
python3 -m py_compile scripts/*.py 2>/dev/null
# Placeholder search
grep -nE '\[TODO\]|\[TBD\]|\[PLACEHOLDER\]|\[INSERT\]' {file}
# Untagged code blocks
grep -c '```$' {file}
Gate: All code checks complete. Proceed only when gate passes.
Phase 5: Integration Verification
Goal: Confirm cross-references and tool declarations are consistent.
Reference Resolution:
- Extract all referenced files from SKILL.md (grep for
references/) - Verify each reference exists on disk
- Check shared pattern links resolve (
../shared-patterns/)
Tool Consistency:
- Parse
allowed-toolsfrom YAML front matter - Scan instructions for tool usage (Read, Write, Edit, Bash, Grep, Glob, Task, WebSearch)
- Flag any tool used in instructions but not declared in
allowed-tools - Flag any tool declared but never used in instructions
Anti-Rationalization Table:
- Check that References section links to
anti-rationalization-core.md - Verify domain-specific anti-rationalization table is present
- Table should have 3-5 rows specific to the skill's domain
# Check referenced files exist
grep -oE 'references/[a-z-]+\.md' skills/{name}/SKILL.md | while read ref; do
ls "skills/{name}/$ref" 2>/dev/null || echo "MISSING: $ref"
done
# Check tool consistency
grep "allowed-tools:" skills/{name}/SKILL.md
grep -oE '(Read|Write|Edit|Bash|Grep|Glob|Task|WebSearch)' skills/{name}/SKILL.md | sort -u
# Check anti-rationalization reference
grep -c "anti-rationalization-core" skills/{name}/SKILL.md
Gate: All integration checks complete. Proceed only when gate passes.
Phase 6: Generate Quality Report
Goal: Compile all findings into the standard report format.
Show all test results with individual scores — never summarize as "all tests pass." Sort findings by impact (HIGH / MEDIUM / LOW). Include specific, actionable recommendations with file paths and line numbers. When batch evaluating, show how each item compares to collection averages; do not report "most are good quality" without quantitative data.
This phase is read-only: report findings but never modify agents or skills. Use skill-creator for fixes. Clean up any intermediate analysis files created during evaluation.
Use the report template from references/report-templates.md. The report MUST include:
- Header: Name, type, date, structural score, maximum, and grade
- Structural Validation: Table with each scorer check, status,
earned/max, and detail - Qualitative Analysis: Evidence-backed findings kept separate from the score
- Code Quality: Script syntax results, placeholder count, untagged block count
- Issues Found: Grouped by HIGH / MEDIUM / LOW priority
- Recommendations: Specific, actionable improvements with file paths and line numbers
- Comparison: Score vs collection average (if batch evaluating)
Issue Priority Classification:
| Priority | Criteria | Examples |
|---|---|---|
| HIGH | Broken functionality or a severe structural failure | Syntax errors, invalid frontmatter, broken critical references |
| MEDIUM | Incomplete or misleading guidance | Stale instructions, weak recovery guidance, tool mismatch |
| LOW | Cosmetic or minor quality issues | Untagged code blocks, missing changelog |
Grade Boundaries (percentage of total / max_total):
| Score | Grade | Interpretation |
|---|---|---|
| 90-100 | A | Strong structural health |
| 75-89 | B | Good structural health |
| 60-74 | C | Structural gaps to address |
| 40-59 | D | Significant structural gaps |
| <40 | F | Major structural gaps |
Gate: Report generated with all sections populated and evidence cited. Evaluation complete.
Examples
Example 1: Single Skill Evaluation
User says: "Evaluate the test-driven-development skill" Actions:
- Confirm
skills/testing/test-driven-development/exists (IDENTIFY) - Run
score-component.pyand record all eight checks (STRUCTURAL) - Inspect content for useful, accurate, proportionate guidance (CONTENT)
- Syntax-check any scripts, find placeholders (CODE)
- Verify all referenced files exist (INTEGRATION)
- Generate scored report (REPORT) Result: Structured report with score, grade, and prioritized findings
Example 2: Collection Batch Evaluation
User says: "Audit all agents and skills" Actions:
- List all agents/.md and skills//SKILL.md (IDENTIFY)
- Run Steps 2-5 for each target (EVALUATE)
- Generate individual reports + collection summary (REPORT) Result: Per-item scores plus distribution, top performers, and improvement areas
Example 3: Structural Compliance Check
User says: "Check the structural health of systematic-refactoring" Actions:
- Confirm
skills/systematic-refactoring/exists (IDENTIFY) - Run the deterministic 90-point precheck (STRUCTURAL)
- Inspect guidance and examples for accuracy and usefulness (CONTENT)
- Run code checks where scripts or examples exist (CODE)
- Generate a report that separates scored checks from qualitative findings (REPORT) Result: Structural score plus evidence-backed qualitative findings
Error Handling
Error: "File Not Found"
Cause: Agent or skill path incorrect, or item was deleted
Solution: Verify path exists with ls before evaluation. If truly missing, exclude from batch and note in report.
Error: "Cannot Parse YAML Front Matter"
Cause: Malformed YAML — missing --- delimiters, bad indentation, or invalid syntax
Solution: Flag as HIGH priority structural failure. Score YAML section as 0/10. Include the specific parse error in the report.
Error: "Python Syntax Error in Script"
Cause: Validation script has syntax issues
Solution: Run python3 -m py_compile and capture the specific error. Score validation script as 0/10. Include error output in report.
Error: "Documented JSON Key Missing"
Cause: The evaluator read internal earned_points or max_points names instead of the JSON contract.
Solution: Read checks[*].earned and checks[*].max; confirm top-level total, max_total, and grade before reporting.
References
Reference Files
${CLAUDE_SKILL_DIR}/references/scoring-rubric.md- Full/partial/no credit breakdowns per rubric category${CLAUDE_SKILL_DIR}/references/report-templates.md- Standard report format templates (single, batch, comparison)${CLAUDE_SKILL_DIR}/references/common-issues.md- Frequently found issues with fix templates${CLAUDE_SKILL_DIR}/references/batch-evaluation.md- Batch evaluation procedures and collection summary format