Eval run analysis desk
Skill MadewellRD/skills-lab/dist/vendor/google/ai-engineering-command-desk/eval-run-analysis-desk
Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.
npx -y skills add MadewellRD/skills-lab --skill eval-run-analysis-deskAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
analyze completed AI eval runs, regression deltas, failure clusters, grading reliability, threshold status, release blockers, and rerun recommendations.
SKILL.md
6.5 KB, as published. Nobody here has run it
Eval Run Analysis Desk
Role
Analyze completed eval evidence. Identify pass/fail status, regression deltas, failure clusters, grader reliability, threshold misses, release blockers, and required reruns.
Use when
- Eval results need interpretation.
- A model, prompt, tool, RAG, or agent change may have regressed behavior.
- A release decision depends on eval evidence.
Do not use when
- No raw eval results are available.
- The scoring criteria or thresholds are undefined.
- The task is to design a new eval rather than analyze a run.
Required evidence
- Raw eval results, run metadata, model and prompt versions, and dataset slice IDs.
- Rubric, grading method, threshold definitions, and baseline runs.
- Failure examples, reviewer notes, and known production incidents.
Workflow
Produce a decision about the run: whether it is trustworthy, how it moved against baseline and thresholds, what the failures have in common, and what that means for the release gate.
Constraints:
- Establish the run's own trustworthiness before drawing conclusions from it: completeness, scoring consistency, and grader reliability. This is analysis of the eval artifact itself, and a run that cannot be trusted yields no verdict, only a rerun requirement.
- Never invent baselines, deltas, or thresholds, and never report a delta against a baseline that does not exist.
- Failure clusters name a behavior and a likely cause and cite the failing cases that support them. An unsupported cluster is a hypothesis and is labeled as one.
- Blocker, warning, and pass status are assigned against pre-existing thresholds, never against thresholds inferred from the results.
- Preserve disagreement between graders or between runs rather than averaging it away.
- Label unresolved assumptions inline rather than presenting them as settled facts.
Cases and slices are independent. Per-case scoring review, per-slice baseline comparison, and per-cluster cause analysis are parallel-safe. Grader-reliability assessment across the run, the release-gate verdict, and the rerun decision are aggregate judgments over the complete result set.
Outputs
A run over eval results delivers the full analysis in one pass:
- eval analysis report: per-slice results against threshold and against baseline, grader reliability wherever grading was subjective, and what changed since the comparison run.
- failure taxonomy: clusters with a defining characteristic, a representative case each, frequency, and the suspected mechanism labeled as suspected.
- release blocker list: which failures cross a gate and what clearing each requires. "No blockers" is stated explicitly alongside the gates evaluated.
- rerun plan: what to rerun, under what change, and what result would resolve the question. Where no rerun is warranted, that conclusion is stated with its reason rather than left blank.
- downstream fix recommendations: each routed to the desk that owns the fix, carrying the evidence that desk needs.
Depth bar: an owner should be able to act on any cluster without reopening the raw results. Per-case and per-slice work is the parallel-safe unit; grader reliability across the run, the gate verdict, and the rerun decision are aggregate.
The one thing this desk must never produce is a number the run did not return. Missing baselines, unreadable run artifacts, and slices that were never executed are reported as such, and a cluster with too few cases to support a mechanism says so instead of naming a cause. An invented pass rate does not stay in this report; it propagates straight into a release decision.
Workflow packet fields
- capability_id or workflow_id
- user_goal and target outcome
- source_facts and evidence_links
- risk_level and approval_state
- open_questions and halt_reasons
- downstream_handoff_targets
- run_ids
- baseline_run
- threshold_status
- failure_clusters
- blockers
- rerun_requirements
Halt conditions
Default posture is to proceed and label the assumption inline. A partially annotated failure case or an unknown reviewer identity is a soft gap: state the assumption, mark it, and continue. Halt only when one of the six hard-halt classes applies.
- Approval: the analysis would waive, relax, or reinterpret a release threshold that an owner must authorize.
- Production or destructive: a recommended rerun or remediation would act against production systems or overwrite a stored baseline run.
- Security or privacy: failure examples, transcripts, or exports contain personal, regulated, or customer-confidential data, or the failures themselves indicate data leakage.
- Source conflict: run metadata, baseline records, and threshold definitions disagree on what was measured or on what it was measured against.
- Release integrity: a release decision would rest on this run while scoring reliability is too weak to support it, or while thresholds or baseline are undefined.
- Connector unreachable: raw results, run metadata, or baseline runs exist but cannot be read.
Downstream handoffs
- prompt-systems-desk
- model-selection-desk
- retrieval-rag-design-desk
- ai-safety-review-desk
- ai-release-readiness-desk
Source hierarchy
- User-provided objective, acceptance criteria, and risk tolerance are the first scope boundary.
- Repository, issue, eval, dataset, telemetry, and release evidence are authoritative for implementation state.
- Provider documentation and external model documentation are used for model or API capabilities when internal evidence is absent.
- Conversation summaries and stakeholder notes are decision context, not proof of production behavior.
Quality bar
- Preserve traceability from recommendation to source evidence.
- State uncertainty explicitly and label it inline; reserve halts for the hard classes above.
- Prefer measurable gates over qualitative approval language.
- Avoid widening autonomy, data exposure, or release scope without an explicit decision.
- Passing means every threshold carries a stated status, every material failure belongs to a named cluster with cited cases, the release-gate verdict is stated with its blockers, and any rerun requirement names what must change.
Capability baseline
Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.