Aipom evaluation strategy advisor
Skill deanpeters/ai-product-operating-model-skills/skills/aipom-evaluation-strategy-advisor
Evidence-based skills for designing and improving AI product operating models
npx -y skills add deanpeters/ai-product-operating-model-skills --skill aipom-evaluation-strategy-advisorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
3 things to look at
- 20 days oldThe repository was created 20 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Recommend the product, model, workflow, human, and production evaluations needed for an AI decision, based on behavior, consequences, evidence gaps, and lifecycle stage.
SKILL.md
5.0 KB, 750 tokens by cl100k_base, as published. Nobody here has run it
AIPOM Evaluation Strategy Advisor
What Is It
Choose the evaluations needed to make a specific AI product decision. Connect behavior and consequence to product, model, workflow, human, and production evidence rather than defaulting to one benchmark.
Why Use It
A model metric can improve while the product decision, workflow, or user outcome gets worse. This advisor makes evaluation coverage and decision rules explicit across the system lifecycle.
When to Use It
Use before major build, launch, expanded autonomy, production review, or after behavior changes and incidents.
What It Produces
- Decision-centered evaluation strategy
- Coverage across behavior, users, workflows, and lifecycle
- Measures, judges, sampling, thresholds, and limitations
- Owners and ship/continue/pause/rollback rules
Who Should Participate
Include the Product Manager, evaluation and technical owners, design or research, operators, affected-user representatives, and governance specialists as consequences require.
Evidence to Bring
Bring the behavior contract, intended-use evidence, real cases, baselines, failures, workflow measures, model results, complaints, monitoring, and current decisions.
How to Do It
- Define the decision evaluation must enable.
- Extract users, behaviors, consequences, lifecycle stage, and existing evidence.
- Map evidence needs across product outcome, system behavior, workflow, human review, and production.
- Identify representative normal, edge, subgroup, adversarial, and severe-failure coverage.
- Choose quantitative and qualitative measures, judges, calibration, sampling, and cadence.
- Set thresholds and critical-failure overrides tied to decisions.
- Name owners, uncertainty, monitoring, and the next evidence gap.
Facilitation Protocol
Support guided, context-dump, and best-guess modes. Ask about the decision before the metric. Present numbered evaluation strategies with coverage and tradeoffs. In best-guess mode label unverified thresholds.
Decision Logic
- Prioritize product evaluation when user and outcome value are uncertain.
- Prioritize behavior evaluation when acceptable output and failure boundaries are unclear.
- Prioritize workflow/human evaluation when review quality, burden, or authority is uncertain.
- Prioritize production evaluation when drift, scale, or real-world interaction dominates.
- Require layered evaluation for consequential systems; no average offsets a critical failure.
Completion Criteria
Finish with the decision, evaluation layers, cases, measures, thresholds, owners, gaps, critical overrides, and next evidence action.
Key Concepts
- Evaluation exists to change a decision.
- Representative coverage matters more than convenient volume.
- Human judges require calibration and limitations.
- Production evidence complements rather than replaces pre-launch evaluation.
Organizational Applications
Use for generated content, recommendations, classifiers, assistants, agents, and AI-assisted internal workflows.
Common Pitfalls
- Beginning with available metrics
- Using vendor benchmarks as product evidence
- Omitting humans and workflows
- Testing only average cases
- Setting thresholds without consequences or owners
- Monitoring without a decision rule
Combine With
Use the behavior contract as input, then build representative datasets and scorecards; use production review after launch.
Assets and Templates
Sources
- NIST, AI Risk Management Framework 1.0, January 26, 2023. Supports mapped, measured, managed, and governed AI risk across the lifecycle. Accessed July 16, 2026.