agentsclimarketplace

Keel

Skill broomva/keel/skills/keel

Measures whether a codebase's verification actually touches the world, or is only checking itself. A grounding ratio for agent-maintained systems.

Install
npx -y skills add broomva/keel --skill keel

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 12 days oldThe repository was created 12 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Measure whether a codebase's verification actually touches the world, or whether it is checking itself. Keel gathers the verification edges it can read in a target (GitHub Actions, CircleCI, GitLab CI, Travis, package scripts, Makefile, Rakefile, pyproject and a dozen tool configs) — and reports the surfaces it recognised but could not parse, so blindness never passes for absence. It classifies each edge as anchored, self-referential, or unknown by asking whether the actor being verified can write to the signal's producer, and reports a grounding ratio. Novel cases are judged by the agent and then crystallized into probes — small reviewable scripts — so repeat shapes get cheaper every run. Then routes each ungrounded check to an anchored signal that already exists in the same graph. Use when: (1) auditing whether AI-generated or agent-maintained work is genuinely verified, (2) assessing how AI-native a codebase actually is — measured by what fraction of its verification the agents cannot author, (3) reviewing a CI/CD pipeline for circular verification, (4) someone claims tests pass and you want to know what that claim rests on, (5) measuring verifier independence. Triggers on keel, grounding ratio, grounded, is this actually verified, who checks the checker, circular verification, self-referential verification, verifier independence.

SKILL.md

15.3 KB, as published. Nobody here has run it

Keel

A ship's keel is the reference line everything else is measured from. It is also what keeps the ship from capsizing — an even keel is a stability property, not a decoration.

Keel measures one thing:

A check is only a check if the signal it reads comes from somewhere the thing being checked cannot write to.

Most verification in an agent-maintained codebase fails this. An LLM reviews what an LLM wrote. A doc is validated against another doc. A status field says "passed" because something set it to "passed". The pipeline is green and nothing has been verified.

The classification

Every verification edge gets exactly one class. The question is never "is this a good check" — it is who produces the signal, and can the actor write to that producer?

ClassMeansExamples
anchoredThe producer is outside the actor's write boundarya test process exit code, a type checker, a payment that settled, a third-party API, a customer action, an independent prober
self_referentialThe producer is inside itan LLM judging output, a doc checked against a doc, a self-set status field, an agent asserting it completed
unknownThe fork point could not be establishedanything you cannot trace
not_a_checkIt asserts nothing about correctnessa dev server, a help target, a formatter that only rewrites, a step that cannot fail (|| true, echo ok)

unknown fails closed — it counts against the ratio, exactly like self_referential. Absence of evidence of dependence is not evidence of independence. And unknown is never settable by the thing being measured: if the target could choose its own class it would never choose unknown.

not_a_check is excluded from the ratio — a node that asserts nothing would be a lie in either column. But it is the one shoppable class: mis-filing a real check here shrinks the denominator and inflates the score. So it carries the same burden of argument as any other verdict, the report prints its count beside the ratio, and the audit samples it like everything else. If you reach for it because a node is hard, the honest answer is unknown.

Grounding ratio = anchored / (anchored + self_referential + unknown).

The loop

Run these in order. Stages 3 and 4 are optional — stage 2 alone produces a complete, honest report.

Run it

The scripts are plumbing — they locate, merge, validate, render and sandbox. Every judgment in the run is yours. Paths are relative to the installed skill; substitute wherever your harness put it.

# 1. locate the surfaces (mechanical, no model call)
bun <skill>/scripts/gather.ts <target> --json > nodes.json \
    --coverage coverage.json          # what it recognised but could not read

# 2. try the probe cache; everything it cannot decide comes back to you
bun <skill>/scripts/classify.ts nodes.json --json > classified.json
bun <skill>/scripts/classify.ts nodes.json --batches   # the same nodes, batched for judging

# 3. YOU judge every pending node over its `raw`, and write Verdict[] to verdicts.json

# 4. merge and validate — refuses on an unjudged node, a naked argument, a bad class
bun <skill>/scripts/assemble.ts nodes.json classified.json verdicts.json \
    --dir <target> --probes <skill>/probes -o report.json

# 5. the artifact
bun <skill>/scripts/render.ts report.json -o report.html

# optional
bun <skill>/scripts/route.ts report.json --dispatch      # then author RouteProposal[]
bun <skill>/scripts/route.ts report.json --proposals p.json -o bindings.json
bun <skill>/scripts/audit.ts report.json --sample 0.1    # the ε-audit; see §4

assemble is the step that keeps the arithmetic honest: it recomputes the grounding ratio from the verdicts rather than trusting any block you hand it, and it refuses rather than quietly computing a ratio over the nodes that happen to have verdicts. A report silently computed over a subset is the exact failure this tool exists to detect.

1. Gather (mechanical)

Find candidate verification edges. This step locates surfaces; it does not judge them. Look at CI workflow definitions, package/task/make scripts, test configuration, review and branch-protection requirements, deploy and promotion conditions, and any wired integrations. Emit Node[] per schemas/keel.ts.

gather reads what it knows how to read. Pass --coverage <file> and it also records the surfaces it recognised and could not parse — a Jenkinsfile, a .buildkite/, a build.gradle. Carry that into the report: a ratio computed over the residue of a repo whose real CI is Jenkins is not a measurement of that repo, and non-coverage that goes unreported is Keel's own shoppable class.

Carry the literal snippet into raw. Downstream reasoning happens over the real text, never over a summary you wrote — a summary is already a judgment.

2. Classify (agentic, cache-first)

For each node:

  1. Try the probe library. Load probes from the shipped probes/ directory and from ~/.config/keel/probes/. Run match(node); on a hit, run assess(node). A non-null result is the verdict — record decidedBy: 'probe' and move on. This costs no tokens.

    Probe code is executed by loading it, so loading and running probes happen only inside a separate child process with a kill-timer held by the parent: a synchronous while(true) cannot be preempted in JS, so an in-process time guard is fiction. A probe that throws, hangs, or exceeds the budget is skipped with a warning and its nodes fall through to your judgment — never fatal to the run.

    That child is sandboxed on macOS only. sandbox-exec exists nowhere else, so on Linux and Windows the child gets a stripped environment and the kill-timer and nothing more: a probe there can write files, reach the network, and spawn processes. Say so if you report on probe provenance, and see SECURITY.md for the enforced-vs-not table. A run that had to ignore part of its own configuration — a KEEL_* variable a dotenv in the target tried to set — records that in warnings; carry it into the report.

  2. On no match, or on abstention, judge it yourself. Read raw. Ask the only question that matters:

    What actually produces this signal, and can the actor being verified write to it?

    Trace the causal path. A test command is anchored because the runtime decides the exit code and no amount of persuasion changes it. An LLM review step is self-referential because the same class of system that produced the work produces the assessment. A deploy check that reads a status field the deployer sets is self-referential no matter how many green checkmarks it renders.

    Fill writeBoundary.argument with the causal path, not a restatement of the class. "Self-referential because it is self-referential" is a failed verdict. If you cannot trace it, say unknown — that is a real answer and it is often the correct one.

See references/grounding-classes.md for worked cases, including the ones that look anchored and are not.

3. Crystallize (optional)

When you judged a node the library could not, and the shape will recur, write a probe to ~/.config/keel/probes/<id>.v<n>.ts implementing the Probe interface. Probes are versioned in the filename because ProbeMeta.version exists and two versions cannot share one path: minting never overwrites, it writes the next version, and the loader takes the highest version per id and warns about the ones it shadowed. Set KEEL_PROBE_DIR to point the library somewhere else.

Rules:

  • A probe abstains (return null) whenever it is unsure. It may never return unknown. Abstention costs a model call; a wrong confident probe costs correctness.
  • Generalize the shape, never the specific repo. match keys on structure.
  • One probe, one shape. A probe matching everything is a rule table wearing a costume, and it will rot.
  • Probes are code so they can be read, diffed, tested, and rejected. Keep them small enough to review in under a minute.

Minted probes live in ~/.config/keel/probes/ so a skill update never destroys them. Contribute one back by copying it into this repo's probes/ and opening a PR — that is how the library compounds across everyone who runs Keel.

4. Audit (optional, and the honest part)

Probes drift. A generalization that over-matches will mis-classify silently forever, and the failure is invisible precisely because it is cheap.

So: sample a fixed fraction of probe-decided nodes — start at ~10% — and re-decide them agentically with the cached verdict hidden from you. Record the comparison in verdict.audit. On disagreement, narrow the probe's match or retire it.

audit.ts is a stepper, for the same reason corpus.ts is: it cannot re-decide anything, because deciding is yours.

bun <skill>/scripts/audit.ts report.json --sample 0.1 --seed 1 -o pending.json
# it prints the sampled nodes WITHOUT their cached class, probe, confidence or
# argument — you judge them cold, and write Verdict[] to redecided.json
bun <skill>/scripts/audit.ts record pending.json redecided.json --report report.json

The blindness is the whole mechanism. If you can see the cached verdict while re-judging, the audit measures your agreement with yourself and nothing else.

The probe library's agreement rate is Keel's own counter-metric. Report it with its denominator — "1.00 over 10 compared nodes" — because a rate over three nodes and a rate over three hundred are different claims. A system that measures groundedness while refusing to measure its own is telling you something.

record reports disagreement and names the probe. It does not retire anything: retiring a probe is a judgment about the world, and that stays with a human.

Output

Write Report (see schemas/keel.ts) as JSON, then render a self-contained HTML report alongside it. The HTML carries the grounding ratio, the node graph with each class, the write-boundary argument for every verdict, and — across runs — the crystallization curve of cost per node as the probe library grows.

The ratio never travels alone. Print the absolute anchored count and the gathered-surface coverage (nodes by kind) beside it, always. A 1.0 over one edge and a 0.7 over fifty are different claims, and a bare ratio rewards deleting checks — the pair is the guard. A target with zero gathered nodes gets an explicit "nothing gathered" state, never a ratio.

Report unknown prominently. It is the most honest number Keel produces.

Modes

Keel is one skill with modes, not a family of skills.

ModeReadsEmits
keel measurea targetReport — the four stages above
keel routea ReportBinding[] — a route from each ungrounded check to an anchored signal already in the graph
keel audita Reportthe probe library's agreement rate, with its denominator
keel constructBinding[]counter-metric pairings, arbitration, audit loops — not yet built
keel applyBinding[]a diff or PR — gated, never completed by an agent

The routing rule

Independence cannot be manufactured, but it can be routed. A route may only point at a node that is present in the same report and already classified anchored. Validate that in code — a proposal that fails the check becomes null with a reason. "No route found" is a first-class answer and is correct whenever the fix needs a policy decision rather than a rewiring.

Route mode runs in three steps: --dispatch emits the judgment payload, the agent authors RouteProposal[] from it, and --proposals validates those into bindings. A proposal may carry an effort, and the only legal values are:

effortmeans
configa value in a file that already exists — a flag, a limit, a needs: edge, an existing job's if:
wiringnew plumbing between things that already exist — a step that reads an artifact another step already produces
processa change to how people or systems behave — a required check, a branch-protection rule, a third party's involvement

--dispatch carries that same list, with the same distinctions, in effortValues — so a proposal can be authored from the payload alone. Effort ranks cheapest-first and never scores; there is deliberately no numeric weight, because a cost is one step from an objective over the ratio. An unrecognised value is warned about and dropped, never coerced to a default — and since a dropped effort is no effort, that route sorts behind every route that stated one. Omitting effort is legal and silent, and lands in the same tier: "we did not say" is not cheap, it is unmeasured. Only an effort the proposal itself declares counts — one inherited from the prototype chain is not something the agent authored, so it is ignored.

Routing never moves the ratio. A proposal is not a change. The number moves only when a human applies one and Keel re-measures from the target — the world stays in the loop, so a route's claim never asserts its own outcome.

The edge that must stay open

The router receives verdicts — what is ungrounded and why. It must never receive the ratio as an objective. The moment the score becomes something to optimize, it becomes a selection signal, and the router hill-climbs into the scorer's blind spot: the number keeps rising while it stops meaning anything.

Inform freely, optimize never.

This is not enforced by putting the router in a separate skill — a package boundary asserts nothing, and would itself be a not_a_check. It is enforced by a test: fabricate an adversarial bindings file claiming everything is routable, re-measure, and assert the verdicts and grounding come back byte-identical.

Scope

Keel measures the shape of verification, not its quality. A repo can be 100% anchored and have terrible tests. Anchoring says the signal comes from outside; it does not say the signal is sufficient. Do not let a high ratio be read as "well tested", and say so in the report.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.