agentsclimarketplace

Run eval harness

Skill urmzd/dotfiles/dot_agents/skills/run-eval-harness

Cross-platform dotfiles managed by Chezmoi with Homebrew/apt and per-language version managers. One-command bootstrap for macOS and Linux with Neovim, Tmux, Zsh, and AI agent skills.

Install
npx -y skills add urmzd/dotfiles --skill run-eval-harness

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Run a project's own eval suite end to end without babysitting: locate the harness, launch it in the background, monitor to completion, parse the report into a normalized summary (pass rate, latency percentiles, tokens, cost), and diff against the previous run with per-case regressions explained. Use when asked to "run the evals", re-run after a change, or score a new dataset in any repo with an eval harness. Do NOT use for the GAP benchmark suite (use run-evals), SAIGE provider-boundary checks (use verify), or designing new eval sets (that is authoring work, not a run).

SKILL.md

3.3 KB, as published. Nobody here has run it

Run Eval Harness

Own the whole run-monitor-parse-diff loop and return once with a complete summary. The user should never have to ask "check the current state".

1. Locate the harness

Search in order; stop at the first hit:

  1. Project docs: README, AGENTS.md, docs/ mentioning "eval".
  2. Task runners: justfile, Makefile, package.json scripts, pyproject scripts with eval targets.
  3. Convention paths: evals/, eval/, benchmarks/, files matching *eval*.py|ts|go.

If multiple harnesses exist, pick the one the user named; otherwise list them and pick the default documented in the repo. Note the dataset in use (golden set path) and where reports land (for example eval_report*, results/, run_log.jsonl).

2. Launch in the background

  • Run via the documented entry point with the repo's own defaults; do not invent flags.
  • Use a background shell so the session stays free; capture stdout to a log file in the scratchpad.
  • Record start time, git commit, model or provider config in effect.

3. Monitor without spamming

Poll the log at an interval matched to expected runtime (a 10 minute run gets checks every 2 to 3 minutes, not every 15 seconds). Detect and report early: crash, auth failure, rate limiting (429 or backoff messages), or a stall with no new output for 3 poll cycles. On transient provider errors, retry the run once before reporting failure.

4. Parse into the normalized summary

Extract whatever subset the report provides:

MetricNotes
Pass / total, pass ratePer strictness level if the harness has them
Latency P50 and P95Report both; never substitute mean
Tokens in / out per caseAnd totals
Cost per run and per caseFrom real configured prices, never hardcoded guesses
FailuresCase id, expected vs actual, one-line cause each

5. Diff against the previous run

Find the most recent prior report or run log. Report: metric deltas, newly failing cases, newly passing cases, and whether config changed between runs (model, dataset, prompt version, commit). If no prior run exists, say so and record this one as the baseline.

6. Report once

Return a single summary: headline result, the metric table, the diff, and failure explanations. State the exact command used, the commit, and the report file path so every number is reproducible. If the harness itself is broken, report the root cause and stop; do not silently fix eval logic, and never edit the golden set or scoring code to make a run pass.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.