agentsclimarketplace

Recipe eval skill

Skill shinpr/rashomon/skills/recipe-eval-skill

Measure prompt and skill improvements with blind A/B comparison.

Install
npx -y skills add shinpr/rashomon --skill recipe-eval-skill

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Creates or updates Claude Code skills through interactive dialog, then evaluates effectiveness with sequential paired comparisons. Use when creating new skills, updating existing skills, or evaluating skill quality.

SKILL.md

5.1 KB, as published. Nobody here has run it

Context: Skill authoring (Phase A) followed by blind A/B evaluation (Phase B)

Mode: $ARGUMENTS

Orchestrator Definition

Core Identity: "I am not a worker. I am an orchestrator."

Execution Method:

  • Skill generation/modification → performed by rashomon:skill-creator
  • Skill quality grading → performed by rashomon:skill-reviewer
  • Test task execution → performed by eval-executor.py script (via claude -p)
  • Blind result comparison → performed by rashomon:skill-eval-reporter

Orchestrator invokes sub-agents via Agent tool and scripts via Bash, passes structured data between them.

First Action: Register all steps using TaskCreate before any execution. Phase A steps are defined in the mode-specific reference (create.md or update.md). Phase B steps are defined in eval.md. Update status using TaskUpdate upon each step completion.

Mode Detection

Determine mode from $ARGUMENTS:

ModeCriteria
Creation"create", new skill request, no existing skill referenced
Update"improve", "update", existing skill name or path mentioned
Unspecified$ARGUMENTS is empty or ambiguous

Scope Boundaries

Phase A (Skill Authoring): Create or modify skill content through dialog. Ends with user-approved skill file. Phase B (Evaluation): Measure skill effectiveness through blind execution comparison. Phase B is read-only for the source skill. A finding that requires authoring changes transitions back to Phase A for user review and approval before evaluation restarts.

Responsibility Boundary: This skill completes with the combined evaluation report and ship/revise/reject recommendation.

Workflow

Phase A: Skill Authoring

Read the mode-specific reference and execute:

Phase A ends with: user-approved skill content (new or modified).

Phase A → Phase B Handoff

Before starting Phase B, confirm these data are available in context. Phase B cannot proceed without them:

DataSourceRequired
Skill namePhase A dialogAlways
Source skill directoryPhase A file writeAlways
Held-out test requestsPhase A Round 3 (create) / Round 2 (update)Always
Trigger scenariosPhase A Round 3 (create) / Round 1-2 (update)Always
Old skill directory snapshot and fingerprintPhase A Step 6 (update mode only)Update mode
Approved source directory fingerprintPhase A final writeAlways

If fewer than two held-out requests are available, ask before proceeding: "What complete requests does your team actually send for work that requires this skill's rules? Please provide at least two verbatim."

Phase B: Evaluation

Read references/eval.md and execute the evaluation protocol. Pass the handoff data above as context.

Phase B consists of:

  1. Trigger check: Does the skill fire for its intended use case? (Step 1)
  2. Trigger fail handling: Diagnose and request an authoring revision when needed (Step 2, conditional)
  3. Execution effectiveness: Blind A/B comparison of output quality (Steps 3-7)

Final Output

Present combined results to user:

  1. Phase A result: Skill quality grade (A/B/C from rashomon:skill-reviewer)
  2. Phase B trigger: Discovered (yes/no), Used (yes/no), usage evidence
  3. Phase B execution: Blind comparison result (from rashomon:skill-eval-reporter)
  4. Recommendation: ship / revise / reject

Error Handling

ScenarioBehavior
User cancels during Phase AStop. No eval needed.
Grade C after 2 iterationsPresent content with issues. User decides: accept/revise/abort.
One executor fails in Phase BPreserve diagnostics, mark comparison inconclusive, and make no winner or effectiveness recommendation.
Both executors fail in Phase BReport failure. Phase A result still valid.
Worktree creation failsReport git error. Phase A result still valid.

Prerequisites

  • Git repository with git worktree lock support
  • claude CLI available in PATH
  • Sufficient disk space for worktree copies

Completion Criteria

Phase A

  • Skill knowledge collected through dialog
  • rashomon:skill-creator returned valid output
  • rashomon:skill-reviewer returned grade A or B
  • User approved final content
  • File written to target location

Phase B

  • Trigger check executed and result presented
  • Sequential paired execution completed in worktrees
  • Blind comparison completed by rashomon:skill-eval-reporter
  • Worktrees cleaned up
  • Combined report presented with recommendation

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.