agentsclimarketplace

Implement and check

Skill skillberry-ai/cap-evolve/skills/phases/implement-and-check

The HARD GATE that must pass before any optimization budget is spent. Use right after intake. Walks the agent through implementing the 3 required adapter methods plus any defaulted hooks that need overriding (and any selected skill's abstract methods), then runs `cap-evolve check` on the project plus each involved skill's check.py, refusing to proceed until everything is implemented and deterministic — and listing exactly what is still stubbed.From its SKILL.md

Install
npx -y skills add skillberry-ai/cap-evolve --skill implement-and-check

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • runs commandsInstructs the agent to run 2 commands, including `python scripts/run.py --project .capevolve/project --skill-check <skills>/capabilities/<cap>/scripts/check.py` and 1 more.

SKILL.md

6.9 KB, ~1.5k tokens by cl100k_base, as published. Nobody here has run it

implement-and-check — make the contract real

Optimizing against a half-wired adapter produces a number that means nothing: if the scorer is a stub, every candidate scores the same; if tasks() is empty, the mean is computed over nothing; if run_target is non-deterministic in a way the scorer can't see, the gate chases ghosts. This phase is the hard gate that ensures the contract holds before a single unit of budget is spent. It is cheaper to fail here than after a full optimization run.

Inputs / outputs (manifest tokens)

  • needs: project — the scaffolded .capevolve/project/ from intake.
  • provides: checked — the proof that the adapter (and any involved skill) is fully implemented and deterministic. baseline will not run without it.

Steps

  1. Implement the 3 required adapter methods in .capevolve/project/adapters/adapter.py (see docs/ADAPTER_CONTRACT.md). These three are @abstractmethod — the gate refuses to run until all three are real:

    • tasks(split) — yield the evaluation tasks (non-empty, stable across calls).
    • run_target(task, ctx, *, seed=0) — run the agent under test with the candidate live as ctx; capture output + trace into a Rollout. Forward seed if the runner is stochastic.
    • score(task, rollout) — return a reward in [0,1] + general feedback (no gold-answer leakage — it becomes the diagnosis signal).

    Then override a defaulted hook only if its default does not fit your capability: materialize(candidate_dir, edits=None) (default: pure write of each {component: text} edit as a file under candidate_dir — override to support a capability-specific patch format), live(candidate_dir) (context manager yielding ctx), apply(candidate_dir, edits=None) (back-compat inject), trajectories(split, ctx=None) / runner_model() (both default to None). Separately, run_batch / run_trials / score_batch are not on the base class — the harness feature-detects them with hasattr and uses them if you define them.

  2. Implement any selected skill's scripts/abstract.py (most are concrete and need nothing).

  3. Run the gate:

    python scripts/run.py --project .capevolve/project \
        --skill-check <skills>/capabilities/<cap>/scripts/check.py
    

    It runs cap-evolve check (adapter: no stubs, tasks non-empty + stable, scorer deterministic, materialize() probed) and each named skill's check.py. Exit 0 = green; the JSON lists exactly what is still stubbed or non-deterministic.

  4. Pipeline-wiring self-test (runs automatically once the check is green). A green adapter is necessary but not sufficient — the optimizer also needs its context wired. After the check passes, run.py runs pipeline_selftest.py, a cheap, benchmark-agnostic check (no API cost) that asserts the plumbing the optimizer depends on:

    • the optimizer-prompt template is scaffolded at optimizer/INSTRUCTIONS.md and still carries its {{...}} placeholders (intake must not delete them);
    • capevolve.yaml::optimizer_instructions_file points at a file that EXISTS;
    • rendering that template through the REAL harness renderer leaves NO {{ placeholder behind (every dynamic block substitutes);
    • the adapter either DEFINES trajectories() (native traj dir → copied verbatim into the optimizer's ./trajectories/) or intentionally inherits the base default (cap-evolve falls back to its per-rollout JSON) — both valid, reported.

    It reports the precise missing/broken artifact so you can iterate until green. (Pass --no-pipeline-selftest to skip it; run it standalone with python scripts/pipeline_selftest.py --project .capevolve/project.)

    A full one-iteration mock run is intentionally NOT done here: it would need a baseline + frozen split + run dir that do not exist yet at gate time, and those are benchmark-specific. This self-test exercises the same workdir-building and prompt-rendering code paths instead — the wiring an optimizer actually consumes.

What the check actually verifies (and why)

  • No stubs — a NotImplementedError/pass body means the method silently returns nothing; the resulting score is meaningless.
  • tasks non-empty and stable — an empty or shuffling task set makes the mean and the split irreproducible.
  • Scorer determinism — score the same rollout twice; differing rewards mean the "reward" includes scorer noise the optimizer cannot learn from. (Target stochasticity is fine and is handled by multi-trial evaluation; scorer nondeterminism is a bug.)
  • materialize() probed — an edit that cannot be materialized cannot be evaluated, so the check calls it against a temp copy (pure, so the host is untouched). This one is a probe, not an assertion: a raise is reported as a note and does not fail the check (core/cap_evolve/check.py:166-167), because a real adapter may need its full environment. Green here means "callable or explained", not "edit path verified".

Do not proceed until green

If it reports stubs or non-determinism, fix them and re-run. This is the standard validation-gate discipline for self-improving systems: prove the measurement apparatus works before you trust any measurement it produces. A green check is the only honest entry into baseline.

What good vs bad looks like

  • Good: {"ok": true} with all 3 required methods concrete, a deterministic scorer, and every involved skill's check green.
  • Bad: proceeding on a red check "to save time"; a scorer that returns different rewards for the same rollout; an empty/placeholder tasks(); feedback that leaks the gold answer (passes the wiring check but corrupts diagnosis).

Dual-mode

This phase runs two ways from the same SKILL.md: standalone as the slash command /cap-evolve:implement-and-check (the argument-hint shows its run.py args), and orchestrator-callable — cap-evolve run / the orchestrate skill invokes the same scripts/run.py headlessly and threads the run dir between phases.

References

  • references/concepts.md — the adapter contract, why each check exists, the scorer-determinism-vs-target-stochasticity distinction, and the validation-gate-before-budget rationale, with sources.

What ships with it: 7 files

16.1 KB alongside SKILL.md, 5 of them executable

references/

scripts/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.