agentsclimarketplace

Eval run analysis desk

Skill MadewellRD/skills-lab/dist/vendor/openai/ai-engineering-command-desk/eval-run-analysis-desk

analyze completed AI eval runs, regression deltas, failure clusters, grading reliability, threshold status, release blockers, and rerun recommendations.From its SKILL.md

Install
npx -y skills add MadewellRD/skills-lab --skill eval-run-analysis-desk

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

6.5 KB, ~1.3k tokens by cl100k_base, as published. Nobody here has run it

Eval Run Analysis Desk

Role

Analyze completed eval evidence. Identify pass/fail status, regression deltas, failure clusters, grader reliability, threshold misses, release blockers, and required reruns.

Use when

  • Eval results need interpretation.
  • A model, prompt, tool, RAG, or agent change may have regressed behavior.
  • A release decision depends on eval evidence.

Do not use when

  • No raw eval results are available.
  • The scoring criteria or thresholds are undefined.
  • The task is to design a new eval rather than analyze a run.

Required evidence

  • Raw eval results, run metadata, model and prompt versions, and dataset slice IDs.
  • Rubric, grading method, threshold definitions, and baseline runs.
  • Failure examples, reviewer notes, and known production incidents.

Workflow

Produce a decision about the run: whether it is trustworthy, how it moved against baseline and thresholds, what the failures have in common, and what that means for the release gate.

Constraints:

  • Establish the run's own trustworthiness before drawing conclusions from it: completeness, scoring consistency, and grader reliability. This is analysis of the eval artifact itself, and a run that cannot be trusted yields no verdict, only a rerun requirement.
  • Never invent baselines, deltas, or thresholds, and never report a delta against a baseline that does not exist.
  • Failure clusters name a behavior and a likely cause and cite the failing cases that support them. An unsupported cluster is a hypothesis and is labeled as one.
  • Blocker, warning, and pass status are assigned against pre-existing thresholds, never against thresholds inferred from the results.
  • Preserve disagreement between graders or between runs rather than averaging it away.
  • Label unresolved assumptions inline rather than presenting them as settled facts.

Cases and slices are independent. Per-case scoring review, per-slice baseline comparison, and per-cluster cause analysis are parallel-safe. Grader-reliability assessment across the run, the release-gate verdict, and the rerun decision are aggregate judgments over the complete result set.

Outputs

A run over eval results delivers the full analysis in one pass:

  • eval analysis report: per-slice results against threshold and against baseline, grader reliability wherever grading was subjective, and what changed since the comparison run.
  • failure taxonomy: clusters with a defining characteristic, a representative case each, frequency, and the suspected mechanism labeled as suspected.
  • release blocker list: which failures cross a gate and what clearing each requires. "No blockers" is stated explicitly alongside the gates evaluated.
  • rerun plan: what to rerun, under what change, and what result would resolve the question. Where no rerun is warranted, that conclusion is stated with its reason rather than left blank.
  • downstream fix recommendations: each routed to the desk that owns the fix, carrying the evidence that desk needs.

Depth bar: an owner should be able to act on any cluster without reopening the raw results. Per-case and per-slice work is the parallel-safe unit; grader reliability across the run, the gate verdict, and the rerun decision are aggregate.

The one thing this desk must never produce is a number the run did not return. Missing baselines, unreadable run artifacts, and slices that were never executed are reported as such, and a cluster with too few cases to support a mechanism says so instead of naming a cause. An invented pass rate does not stay in this report; it propagates straight into a release decision.

Workflow packet fields

  • capability_id or workflow_id
  • user_goal and target outcome
  • source_facts and evidence_links
  • risk_level and approval_state
  • open_questions and halt_reasons
  • downstream_handoff_targets
  • run_ids
  • baseline_run
  • threshold_status
  • failure_clusters
  • blockers
  • rerun_requirements

Halt conditions

Default posture is to proceed and label the assumption inline. A partially annotated failure case or an unknown reviewer identity is a soft gap: state the assumption, mark it, and continue. Halt only when one of the six hard-halt classes applies.

  • Approval: the analysis would waive, relax, or reinterpret a release threshold that an owner must authorize.
  • Production or destructive: a recommended rerun or remediation would act against production systems or overwrite a stored baseline run.
  • Security or privacy: failure examples, transcripts, or exports contain personal, regulated, or customer-confidential data, or the failures themselves indicate data leakage.
  • Source conflict: run metadata, baseline records, and threshold definitions disagree on what was measured or on what it was measured against.
  • Release integrity: a release decision would rest on this run while scoring reliability is too weak to support it, or while thresholds or baseline are undefined.
  • Connector unreachable: raw results, run metadata, or baseline runs exist but cannot be read.

Downstream handoffs

  • prompt-systems-desk
  • model-selection-desk
  • retrieval-rag-design-desk
  • ai-safety-review-desk
  • ai-release-readiness-desk

Source hierarchy

  • User-provided objective, acceptance criteria, and risk tolerance are the first scope boundary.
  • Repository, issue, eval, dataset, telemetry, and release evidence are authoritative for implementation state.
  • Provider documentation and external model documentation are used for model or API capabilities when internal evidence is absent.
  • Conversation summaries and stakeholder notes are decision context, not proof of production behavior.

Quality bar

  • Preserve traceability from recommendation to source evidence.
  • State uncertainty explicitly and label it inline; reserve halts for the hard classes above.
  • Prefer measurable gates over qualitative approval language.
  • Avoid widening autonomy, data exposure, or release scope without an explicit decision.
  • Passing means every threshold carries a stated status, every material failure belongs to a named cluster with cited cases, the release-gate verdict is stated with its blockers, and any rerun requirement names what must change.

Capability baseline

Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.

What ships with it: 3 files

6.2 KB alongside SKILL.md

agents/

assets/

references/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.