Os eval runner
Skill richfrem/agent-plugins-skills/plugins/agent-agentic-os/skills/os-eval-runner
repo for reusable plugins and skills
npx -y skills add richfrem/agent-plugins-skills --skill os-eval-runnerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Stateless evaluation engine that scores and gates skill improvement iterations using headless Python evaluation scripts. Use when the user says "evaluate this skill", "run autoresearch loop on", "optimize this skill", "run the eval loop", or when another agent proposes a change and needs validation.
SKILL.md
3.9 KB, as published. Nobody here has run it
Skill Improvement Evaluator
Stateless evaluation engine that scores and gates skill improvement iterations using headless Python evaluation scripts.
Ownership Boundary (Critical)
What os-eval-runner owns (permanent, version-controlled with this skill)
- Scoring scripts:
./scripts/evaluate.py,./scripts/eval_runner.py - Scaffold script:
./scripts/init_autoresearch.py - Templates:
./assets/templates/autoresearch/(program, evals, results, proposer prompt)
What lives with the target (deployed per experiment)
All experiment state deploys alongside the target (e.g. <experiment-dir>/references/program.md, <experiment-dir>/evals/evals.json, <experiment-dir>/evals/results.tsv). You MUST read the spec from <experiment-dir>/references/program.md and NOT fall back to engine-local config templates.
Phase 0: Intake Interview
Run this interview before starting any loop or evaluation. If enough information is provided in the initial prompt, skip the redundant questions.
- Q1 — What target skill are you evaluating? (Provide path to skill folder)
- Q2 — Where should the experiment files live? (Defaults to target skill directory)
- Q2b — What metric are you optimizing? (quality_score, f1, precision, recall, or heuristic)
- Q3 — What mode? (Loop mode for autonomous improvement vs QA mode for single diff validation)
- Q4 — (Loop mode) How many iterations? (Default: NEVER STOP)
- Q5 — Does evals.json exist? (If missing, scaffold from template)
- Q6 — Does program.md exist? (If missing, scaffold from template)
- Q7 — Does a baseline score exist? (If missing, run evaluate.py with
--baseline)
Two Modes: Summarized
- Mode 1: Autoresearch Loop: Autonomous iterative improvement. The agent identifies failure types, requests mutations via external proposer CLI (Copilot/Gemini), and runs the eval gate iteratively until the budget or target score is met.
- Mode 2: Single-shot QA: Simple gate validation. Evaluates one specific proposed diff against the baseline and decides KEEP (exit 0) or DISCARD (revert, exit 1).
Stage Pointers & Reference Protocols
- Setup: Start a New Experiment — 4-step setup and re-baselining procedure.
- Mode 1: Autoresearch Loop Protocol — Proposer cycles, prompt mutations, and evaluation loop.
- Mode 2: Single-shot QA Protocol — Context acquisition, reverts, and reporting.
- Phase 2b: Overfitting Gate — Holdout set overfitting checks and forced discard logic.
- Phase 5: Self-Assessment Survey — Mandatory evaluator survey guidelines.
Smoke Test & Gotchas
Smoke Test
- Scaffold an experiment:
python3 ./scripts/init_autoresearch.py --experiment-dir temp/test-exp --mutation-target SKILL.md. - Establish baseline:
python3 ./scripts/evaluate.py --skill temp/test-exp --baseline --desc "smoke test". - Validate exit code: Assert
results.tsvis created, and runningevaluate.pyreturns 0.
Gotchas
- Subjective Simulation: Avoid "mentally simulating" routing accuracy. Subjective audits are strictly banned; run Python evaluation scripts.
- Missing Holdout: Starting loops without holdout prompts. This bypasses the overfitting gate, rendering the results invalid.
- Keywords Footgun: Adding too many triggers to frontmatter. This dilutes semantic discrimination and degrades overall router precision.