agentsclimarketplace

Eval refine

Skill pagoda111king/claude-code-pack/template/.claude/skills/eval-refine

Eval rubric refinement engine · Solo Founder edition harness-optimizer. Scans all historical eval results (idea-eval / skill eval / agent eval), identifies dimensions where rubrics are too loose or too strict, and drafts adjustment PR drafts to candidates/. Use when founder says "the evaluation feels off" or "5/6 all 100% PASS — is something wrong?"From its SKILL.md

Install
npx -y skills add pagoda111king/claude-code-pack --skill eval-refine

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

5.5 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it

Eval Rubric Refiner · Evaluation Standard Tuning

When to Trigger

Founder says:

  • "5/6 skills all 100% PASS · is something wrong?"
  • "I feel this idea scored 22 but it's not worth that much"
  • "Post-launch customer feedback doesn't match the eval"
  • Every Sunday during /learn-review run

Or dashboard pipeline Stage 8 triggers.

Input

  • docs/ideas/*/eval.md (idea-eval historical 6-dimension scores)
  • tests/{skills,agents}/results/*.md (skill / agent eval historical pass rate)
  • Actual feedback data (pilot-tracking.md / customer replies / sales conversion)

Output

docs/learnings/candidates/<YYYY-MM-DD>-rubric-refine.md, structured as follows:

---
type: rubric-refine
generatedAt: 2026-04-25
generatedBy: [email protected]
scope:
  - idea-eval
  - skill:lead-intake
  - skill:idea-eval
analyzedRecords: 14
candidatePatches: 3
status: candidate            # candidate / approved / rejected
---

# Rubric Adjustment Candidates · 2026-04-25

## Candidate 1 · idea-eval "Ability to Pay" dimension too loose

**Evidence**:
- 14 ideas · 12 scored ≥4 on this dimension · but only 3 actually generated revenue within 1 year
- Inference: scores didn't distinguish "can pay" from "will pay"

**Suggested Change**:
- Modify prompt to add: "Evaluate not just **ability to pay**, but **probability of paying this year** · ≥4 requires evidence of 'already spending on similar products'"
- Add to rubric should_have: "Interview Q3 must confirm 'what tools are you currently using' with specific numbers"

**Impact**: Re-evaluate all 14 ideas · estimated 4 will drop from 4 to 3 · total score change ≤ 4 points

**Atomic Check**:
- trigger: idea-eval runs "ability to pay" dimension
- action: add should_have check · add prompt wording
- domain: skills/idea-eval
- confidence: 0.7 (based on real 14 data points)

---

## Candidate 2 · skill eval all 100% PASS

**Evidence**:
- 5/6 skills first run 100% · only lead-intake at 75%
- Real problem exposed by lead-intake (Q10 signal strength) drove v1.1 upgrade
- Inference: the other 5 skills' rubric should_have is too broad

**Suggested Change**:
For each 100% skill, add 1-2 picky should_have checks in yaml. Example:
- proposal-gen add: "Quotation section must show 3 tiers (basic/standard/flagship)"
- content-blitz add: "30 pieces of content must maintain brand voice consistency · sample 5 must 100% pass brand-voice check"

**Impact**: Retest 5 skills · estimated 2-3 will land in 75-90% range (healthy signal)

**Atomic Check**:
- trigger: any skill eval pass rate = 100% for ≥2 consecutive runs
- action: add picky should_have (not new features · stricter checks)
- domain: tests/skills/*.yaml
- confidence: 0.85 (based on real 5/6 lesson)

---

## Candidate 3 · ...

Core Rules

  1. Every candidate must have an "Evidence" section · no real data, don't propose
  2. Every change must have an "Impact" estimate · re-evaluation cost / score change range
  3. Every candidate must pass the atomic 4-field check · not atomic, don't accept
  4. confidence < 0.5 → reject immediately (insufficient data)
  5. Same dimension changed 2 weeks in a row = red flag · likely overfitting · pause for 1 week

Decision Flow

Scan historical eval → find anomaly patterns → draft candidate
   ↓
Candidate goes into docs/learnings/candidates/
   ↓
/learn-review weekly review · founder decides promote / skip / reject
   ↓
promote → update corresponding SKILL.md / agent.md / yaml
   ↓
Retest affected cases · verify change is effective

Solo Founder Specific Constraints

  • Promote at most 1-2 per week · more leads to overfitting
  • Every change must be retested · no retest = changed to "look pickier"
  • Keep pre-change rubric in git history · allows rollback

Collaboration with Other Skills

  • /eval runs automatically trigger this skill (hook trigger)
  • /learn-review calls this skill first on Sunday to collect candidates
  • After promote, update corresponding skill / agent using skill-create tool

Prohibited

  • ❌ "I think" changes without evidence (must have real case support)
  • ❌ Changing rubric without retesting (self-deception)
  • ❌ Changing ≥3 rubrics at once (too many variables · can't attribute)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.