agentsclimarketplace

Aipom eval scorecard builder

Skill deanpeters/ai-product-operating-model-skills/skills/aipom-eval-scorecard-builder

Define calibrated AI evaluation metrics, rubrics, judges, thresholds, sampling, uncertainty, ownership, and decision rules tied to behavior and consequences.From its SKILL.md

Install
npx -y skills add deanpeters/ai-product-operating-model-skills --skill aipom-eval-scorecard-builder

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

4.9 KB, 727 tokens by cl100k_base, as published. Nobody here has run it

AIPOM Eval Scorecard Builder

What Is It

Define how AI behavior will be measured and converted into product decisions: metrics, rubrics, judges, calibration, thresholds, sampling, uncertainty, subgroup views, critical failures, ownership, and action rules.

Why Use It

Metric collections fail when they lack representative cases, calibrated judgment, consequences, or decision thresholds. A scorecard makes clear what a passing score permits—and what a critical failure blocks regardless of averages.

When to Use It

Use after behavior and representative cases are defined, before validation, launch, scale, or recurring production review. Recalibrate when behavior, population, context, judges, or consequences change.

What It Produces

  • Metric and rubric definitions
  • Judge, calibration, sampling, and uncertainty plan
  • Thresholds, critical-failure rules, and subgroup views
  • Continue, revise, constrain, rollback, or stop decision rules

Who Should Participate

Include product and evaluation owners, domain experts, data and engineering, affected-user perspectives, and governance partners where consequences are material.

Evidence to Bring

Bring behavior contracts, evaluation strategy, governed cases, baselines, incidents, consequences, subgroup needs, reviewer guidance, judge comparisons, and candidate thresholds.

How to Do It

  1. Define the product decision and behavior dimensions being evaluated.
  2. Select metrics and rubrics that distinguish quality, safety, workflow, human, and outcome evidence.
  3. Define unit, denominator, direction, aggregation, subgroup, and uncertainty for every measure.
  4. Choose human, automated, model-based, or hybrid judges according to consequence and explainability needs.
  5. Calibrate judges against qualified reference decisions and measure disagreement.
  6. Define sampling across representative, edge, adversarial, subgroup, and production cases.
  7. Set evidence-based thresholds and non-compensable critical failures.
  8. Map results to continue, revise, constrain, rollback, or stop actions.
  9. Assign measurement, review, decision, exception, and recalibration ownership.

Key Concepts

  • A metric without a decision rule is observation, not governance.
  • Aggregate performance can hide subgroup or critical failures.
  • Model judges require calibration and monitoring.
  • Thresholds express consequence and tolerance, not universal truth.

Organizational Applications

Use for pre-release evaluation, regression tests, vendor comparison, human-review quality, workflow adoption, production monitoring, and rollback decisions.

Common Pitfalls

  • Choosing metrics because tools expose them
  • Using averages to hide critical failures
  • Treating model judges as objective
  • Setting thresholds after seeing desired results
  • Omitting sampling and uncertainty
  • Assigning measurement without decision authority

Combine With

Use aipom-evaluation-strategy-advisor for coverage, aipom-production-evidence-review for recurring decisions, and aipom-risk-control-incident-playbook for threshold-triggered response.

Assets and Templates

Sources

What ships with it: 3 files

1.5 KB alongside SKILL.md

Gives 0 of the 12 instructions most evals benchmarks skills give in 727 tokens

Counted across 499 of the 513 authors here whose files we hold, read 2026-09-06

  • Spawn with-skill and baseline runs in the same turnin 31 of 499, across 24 files
  • Keep SKILL.md under 500 linesin 31 of 499, across 24 files
  • Draft assertions while test runs are in progressin 31 of 499, across 24 files
  • Compare against the baseline after changesin 31 of 499, across 13 files
  • Define evals before codingin 26 of 499, across 17 files
  • Run evals frequently during developmentin 25 of 499, across 16 files
  • Keep evals fastin 24 of 499, across 15 files
  • Version evals with codein 24 of 499, across 15 files
  • Generate the eval viewer before evaluating outputs yourselfin 24 of 499, across 17 files
  • Generate an eval report after runsin 24 of 499, across 15 files
  • Track pass@k metrics over timein 22 of 499, across 14 files
  • Save a baseline before making changesin 21 of 499, across 9 files

Said here and by no other author read

  • Define the product decision and behavior dimensions
  • Select metrics distinguishing quality, safety, workflow, human, outcome evidence
  • Define unit, denominator, direction, aggregation, subgroup, uncertainty per measure
  • Choose judges matching consequence and explainability needs
  • Calibrate judges against qualified reference decisions and measure disagreement
  • Define sampling across representative, edge, adversarial, subgroup, production cases

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.