Generate tests
Skill livlign/claude-skills/plugins/dotnet-coverage-kit/skills/generate-tests
A Claude Code plugin marketplace for open-source maintainers — README heroes, README audits, and .NET test-coverage backfills.
npx -y skills add livlign/claude-skills --skill generate-testsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 18 stars18 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Generate unit tests for a .NET service following the coverage-kit conventions. Use when the user asks to 'generate tests', 'backfill tests', 'characterize this service', or 'add tests for' a target after coverage-init has run. Operates in characterization mode (freeze current behavior for existing code) or spec mode (assert intended behavior for new/changed code), enforces the run-capture-fill loop, routes untestable units to the manifest cannot-test log, can fan the backfill out across a user-chosen number of parallel agents for large worklists (asking first, since more agents cost more tokens), and runs a single read-only suite critique before the baseline locks in (auditing cannot-test legitimacy, assertion/C1 depth, and systematic gaps — never the C0/C1 numbers themselves).
SKILL.md
28.0 KB, as published. Nobody here has run it
generate-tests
Generates tests under the base rules and this repo's overlay + manifest. Read all three before generating:
${CLAUDE_PLUGIN_ROOT}/rules/unit-testing.base.md.claude/coverage/refs/unit-testing.md(repo overlay).claude/coverage/refs/coverage-manifest.yml(categories, exclusions, gate)
Build the full worklist first — then cover ALL of it (do not stop early)
The scope of a backfill pass is the entire testable set the manifest defines, not a handful of "priority" files. Before writing any test, enumerate that set explicitly from the manifest and write it down as a checklist. The testable set is:
- Every file matched by a
category_maptarget category and not matched by anexclusionspattern (the genuine unit-scope files), AND - Every carve-out method on an exclusion entry: the structured
carve_outs:list (each{ method, lines, testable }), or the legacyCARVE-OUT: MethodA, MethodBprose inreasonon older manifests. These are the thin pure/decision slices of god-classes (the UserService pattern): testable units even though their file is integration-scope, and a DbContext-injected service is tested by backing the context with the EF in-memory / SQLite provider. Init recorded them precisely so this step covers them; the file'sexcluded_restsays what stays out of scope.
Resolve every worklist item to a full repo-relative path, not a basename. Repos have many
files sharing a name (UserService.cs, Handler.cs); test the exact file the manifest entry
points at, mirror the test under that same path, and never carry one file's carve-out methods onto
its same-named sibling. If a manifest pattern is ambiguous (bare basename matching several files),
flag it back to the manifest rather than guessing which file it meant.
Items already covered by existing tests are checked off, not re-done. Items in cannot_test
are out. Everything else in the set MUST be either tested in this backfill or appended to
cannot_test with a reason — there is no third option of silently leaving it untouched.
An item is "done" only when its branches are covered — not when it has a green test. This
is the definition that matters and the one most often shortcut. A method with filters, a
switch, asc/desc, null guards or error paths needs an input pinned per branch; one happy-path
test checks the item off the worklist while leaving 40–50% of its branches cold. A test
existing is not the bar — its branches being exercised (or sent to cannot_test with a
reason) is. Treat an under-covered target method exactly like an untested worklist item: not
done. See "Cover the BRANCHES" below for how.
Hard rule: a coverage percentage is never a stopping condition — in either direction. Do
not stop because the number "looks low" or "looks done," and do not stop at worklist
completeness while target-layer methods still have uncovered branches. The pass ends only when
every worklist item has branch-covered tests or a cannot_test entry. The classic failure
this prevents: writing one test per file, watching the suite go green, and stopping at ~50–60%
C0/C1 because each method got one path — that is an unfinished pass, not a complete one.
Phasing — when the worklist is large, plan it and report progress
If the worklist is large enough that one pass is impractical (rule of thumb: more than ~15 target files or ~40 carve-out methods, or it spans multiple test projects), do not silently do a slice and stop. Instead:
- Split the worklist into phases (group by file/service/test-project so each phase is a coherent, runnable chunk).
- Tell the user the phase plan up front — the full count of testable units, how many phases, and what each phase covers — before generating. This is the "let me know if it needs phases" contract.
- Execute phases in order. After each phase, report progress as N of M worklist items done and what remains. Keep going through all phases unless the user says otherwise.
- The working tree stays dirty across phases; do NOT commit between phases.
Scaling the backfill — efficiency, then optional parallel fan-out
A legacy worklist can be enormous. Two levers: make each unit cheaper, and (optionally) do many at once. The efficiency rules apply whichever way you run — sequential or parallel:
-
Batch run-capture-fill per class, not per method. The dominant cost is build + test execution. Write the whole test class for a file — one test per branch of each method, not one test per method — build once, run the class once, capture all actual values from that single run, then fill. The economy is in batching the build/run, not in writing fewer tests: never pay a build/run cycle per method, and never trade away branch coverage to save cycles.
-
Pre-triage narrows the assertion, it does not discard the unit. Before the write-build-run loop, scan each unit for an untestable signal (
DateTime.Now/UtcNow,Guid.NewGuid,Random, direct infra/IO). A signal is a reason to look, not a verdict — a whole method iscannot_testonly when the untestable condition is confirmed, not merely present:- Confirm "no seam." Is the dependency really un-mockable? A collaborator injected as an
interface/abstract, an
IHttpClientFactory, an optionalFunc<DateTime>— these are seams. A constructor-injectedDbContextis also a seam: back it with the EF in-memory / SQLite provider. ADbContexton the constructor is NOT grounds to blanket-cannot_testthe class; dumping every DbContext-touching method intocannot_test(as a real run did, roughly 87 false entries in one repo) is the canonical false-exclusion this rule exists to stop. Route tocannot_testonly when the nondeterministic/infra call is made directly with no injectable seam, or when the specific EF op is one the in-memory provider cannot run (raw SQL,ExecuteUpdate/ExecuteDeleteAsync,Batch*, real transactions, unset-rowversion inserts, untranslatable LINQ). - Confirm the nondeterminism reaches the assertion. A method that logs
DateTime.Nowbut returns a deterministic result is fully testable — assert the result. Only when the nondeterministic value flows into the output you would assert (and cannot be pinned) does that assertion drop. Named anti-pattern — the entry-lineStopwatch/Guid.NewGuid: event handlers and command handlers commonly open withvar sw = new Stopwatch()/Guid.NewGuid()for timing and a correlation id used only in a log line. That is NOT the method's behavior — the behavior is the event→document/state mapping that follows. Do NOT route the whole handler tocannot_testfor it; test the mapping/branches and ignore the log. Routing ~50 near-identical handlers tocannot_teston this signal (as a real run did) is the canonical false-exclusion this rule exists to stop. - Still cover the deterministic branches. A method with a
Guid.NewGuid()id and three validation branches gets tests for the three branches (assert everything except the id, or assert the id is non-empty).cannot_testscopes to the specific method/assertion that is genuinely blocked — never the surrounding testable logic.
When in doubt, run the loop — let run-capture-fill prove testability rather than pre-declaring it away to save a cycle. Every
cannot_testentry you do record carries a mitigation (the seam that would unlock it), per the base rule. - Confirm "no seam." Is the dependency really un-mockable? A collaborator injected as an
interface/abstract, an
-
Pattern-replicate across siblings. Legacy services are full of structurally identical units (handlers, validators, mappers). Solve the seam/mock setup once on a representative, then replicate the shape across the lookalikes. A "phase 0" that builds the shared fixtures/builders once feeds this.
-
Risk-order the worklist by complexity × churn × fan-in so the most valuable coverage lands first and the hardest seam problems surface early (while there is budget to fix them). This does not reduce total work — the worklist is still covered in full — it just sequences it well.
Optional parallel fan-out (ask the user first)
The worklist is independent per file, so it parallelizes cleanly. For a large worklist, offer to run the backfill as a fan-out workflow — but never auto-spawn agents. Parallel agents burn tokens fast, and the right count depends on how quickly the user wants it done and their budget, so present the trade-off and let them choose:
-
Build and risk-order the full worklist (above). State the total count of testable units.
-
Ask the user how many agents to run, mapping count to a rough wall-clock for this worklist (scale the examples to the real size; caveat that they are estimates and that more agents = proportionally more token burn):
- 1 agent — sequential, lowest token burn, ~a day for a big service.
- 3 agents (default) — balanced, ~6 hours.
- 10 agents — fastest, highest burn, ~1 hour.
-
On their choice, write the risk-ordered worklist to disk as a JSON array at
coverage/backfill/worklist.json(each item shaped{ file, methods?, group?, mode? }), then invoke the workflow with the manifest path, not the list itself. Get the call exactly right (each rule below is a real field failure this kit has hit, where ZERO agents ran and the launch still looked like success):- Write the file FIRST and confirm it is non-empty (
test -s coverage/backfill/worklist.json). The workflow's Plan agent reads this file off disk; if it is missing or empty the run bails withempty-worklistand spawns nothing. - Pass
argsas a real JSON object, never a JSON-encoded string. A stringified payload makes every field undefined inside the workflow, so it silently defaults and runs nothing. - Pass
worklistManifest(the path). There is NOworklistarg. An inlineworklist:[...]is ignored by the workflow AND, when large, trips the approval-dialog control-char guard. Do not put the array inargsat all. - Invoke via
scriptPath, notname(name + inline args is what tripped the guard in the field).
Workflow({ scriptPath: "${CLAUDE_PLUGIN_ROOT}/workflows/coverage-backfill.workflow.js", args: { concurrency: <chosen>, solution: "<sln>", worklistManifest: "coverage/backfill/worklist.json" } }). If the user just says "go", defaultconcurrency: 3. - Write the file FIRST and confirm it is non-empty (
-
After launch, drive it to completion. Do not re-ask, re-explain, or fall back to one-by-one. Once concurrency is chosen and the workflow is launched, the fan-out IS the plan; executing it is the goal, not a checkpoint. Monitor the background run to done. If a run dies mid-flight (session exit, interruption), resume it with
resumeFromRunIdso completed chunks replay from cache; do NOT fresh-relaunch, which orphans the prior run's git worktrees and duplicates work. Only start a truly new run when there is no prior run to resume. Never quietly abandon the workflow and generate the remaining worklist with sequentialAgent()calls: that is the exact regression this fan-out exists to prevent (it burns the same tokens without the parallelism and drags the user through dozens of turns). If the workflow is genuinely unusable, say so and ask; do not silently downgrade.
The workflow reads the worklist off disk and partitions it into concurrency chunks (one agent
each, in its own git worktree so parallel writes don't collide), generates + adversarially verifies
each chunk, then a
single assemble step applies every chunk's patch to the main tree, merges cannot_test discoveries
once (no manifest write races), and confirms the full suite is green. It leaves the tree dirty:
the promotion gate below still governs when the baseline report and commit happen.
Measure for feedback DURING the pass (this is not the baseline)
You cannot judge branch completeness by eye — measure it, and measure it before you decide you are done, not after. Measuring for feedback and locking the baseline are two different acts; only the latter is a promotion step. So:
- In-loop, as often as useful: run
./.claude/coverage/tools/report.shread-only to see C0/C1 and the Risk Hotspots table. This writes a throwawayREPORT.mdbut you do NOT writebaseline.recorded_overallfrom it — it is a feedback instrument, nothing more. - Read the signal and act: every target-layer (non-
excl:) row in Risk Hotspots, and every target method with an uncovered branch that is not incannot_test, is an unfinished worklist item. Go back and pin inputs for its missing branches (run-capture-fill). Re-measure. Loop until every target-layer method is either fully branch-covered or its remaining uncovered branches are logged incannot_test— i.e. 100% of the testable set, per the exit gate. (The manifest's diff-coverage branch threshold is a per-PR floor, not the backfill target — do not stop the backfill at it.) This is the part that moves a pass from ~50–60% to full testable coverage — and it belongs in the generation loop, not in a later cleanup.
This is explicitly NOT "stopping on a number" — you are not chasing a percentage, you are discharging worklist items whose branches the number reveals are still uncovered. The pass is done when that list is empty, whatever the resulting percentage happens to be.
Exit gate — the testable set is 100% branch-covered; the only residual is documented cannot_test
The backfill exits ONLY when the entire testable set is covered — not "most", not "the priority
files", 100% of it. The testable set is: every target unit-scope method + every carve-out method
(the pure/branching slices of excluded files, including controllers and IO orchestrators), with
all branches exercised (see "Cover the BRANCHES"). The single permitted residual is
cannot_test, and it is not a free pass:
- Every uncovered branch is either covered or a
cannot_testentry — there is no third state. If a target-layer method sits below 100% branch coverage and it is not incannot_test, the pass is not done; go pin the missing branches. A worklist item is closed only when its branches are exercised OR the specific blocked branch/method is incannot_test. - Each
cannot_testentry is a documented, mitigated decision, scoped to a method (never a file/folder):{ target, category, lines, reason (signal at file:line), mitigation }. The mitigation is the concrete source change that would unlock it (extractIClock, inject the repository interface, move the pure slice out), or an explicit "none — genuinely nondeterministic external boundary". This is the "detailed report + mitigation plan for what truly cannot be covered" that the exit requires. - The residual is reported, not buried. Report §6 ("Not Testable") renders every entry with
its lines, reason, and mitigation; the suite critique (below) audits each one hardest before
the baseline locks. A large
cannot_testlist is a signal to escalate a shared-seam refactor to the human, not to accept as-is.
This is not "stopping on a number" — you are not chasing the overall %, you are driving the testable set's uncovered branches to zero. The resulting overall % is whatever falls out of that.
The manifest target (default C0 95% / C1 85% on the Adjusted slice) is an ACCEPTANCE bar, not
this exit condition. The exit gate above is stricter: it drives the testable set to 100% branch
coverage with cannot_test as the only residual. The target is a downward sanity check on the
result: once the exit gate is met, if the measured Adjusted is still BELOW target, the residual
(the cannot_test list / exclusions) is too large to be honest, which is a signal to re-open the
sweep or escalate a shared-seam refactor (the systematic-seam clusters in the suite critique), not
to lower the bar. Do not confuse the two: you never stop early at the target, and you never accept
an Adjusted below it without escalating why. The target is also what the diff gate is set to, so new
code lands at the goal; the recorded baseline floor sits a few points below it (see the manifest).
Promotion gate — set the baseline and commit only when the exit gate is met
coverage-report (to set/raise the baseline) and any commit are promotion steps. They
run ONLY after the exit gate above is met (the testable set 100% branch-covered, every residual a
mitigated cannot_test entry) and the suite critique below has run. Per
the action ladder, generation itself is edit-only: leave the tree dirty and stop. The
distinction from the in-loop measurement: that one is read-only feedback you may run anytime;
this one writes the floor, so it runs once, last. Do not write baseline.recorded_overall
mid-backfill, and do not propose committing a partial backfill — a partial baseline locks in a
misleadingly low floor.
Measure via the wrapper, not by hand. When you do measure, run the one-command wrapper
./.claude/coverage/tools/report.sh (it collects with the repo filefilter, joins the manifest, and
persists a dated reports/<YYYY-MM-DD>/REPORT.{md,html} + CANNOT-TEST.md). Do NOT hand-run run-coverage.sh + coverage-gate.py and
read the numbers off stdout — that leaves no saved report and skips the repo scoping, so the
report effectively goes missing. If report.sh isn't installed yet (repo pre-dates it), copy
it from ${CLAUDE_PLUGIN_ROOT}/scripts/report.sh first, then run it. Only after REPORT.md
exists do you write the measured Adjusted into baseline.recorded_overall.
Never record a baseline off a red suite. Coverage measured while any test fails is
unreliable — a failing test may not have executed the lines it was meant to, so the numbers are
not the real output. If report.sh's Test Results show any failure (the report prints a ⚠️
banner), the baseline step is blocked: fix every failure until the suite is fully green, then
re-measure. A backfill that ends with, say, 54 failing tests is unfinished, not ready to
promote — do not write baseline.recorded_overall from that run.
Suite critique — once, before the baseline locks in
When the worklist is exhausted and the suite is green, run a single critique over the finished
suite before coverage-report sets the baseline. This is the natural moment: fixing is
cheapest now, and a dodged cannot_test or shallow tests would otherwise freeze a misleading
floor. Run report.sh first so the critique can use the report's analytics (e.g. low-C1
buckets) as leads.
Spawn one read-only subagent over the whole suite + the manifest + the report. Single and read-only on purpose (same reasoning as the init critique): systematic problems — a bucket dodged wholesale, the same smell everywhere — are only visible to one reviewer seeing all of it, applying one standard; it produces findings, it does not silently rewrite tests. Its mandate:
cannot_testlegitimacy. Re-check everycannot_testentry against the rubric signal: is it genuinely untestable (no seam, real infra, nondeterministic), or did generation take the easy out? Flag any entry where a seam exists or a deterministic test is feasible. This is the escape hatch — audit it hardest. Watch specifically for the entry-lineStopwatch/Guid.NewGuidfalse-exclusion (a whole handler dropped because it opens with a timing/correlation id used only in logging) — the handler body is testable and must be covered.- Assertion quality / C1 depth. Find files where C0 is high but C1 is low (lines run but branches never asserted → shallow tests) and any vacuous/tautological assertions the per-chunk verify missed (asserting a mock returns what it was set to). Whole-suite view — do not re-do the per-test check the backfill workflow already did.
- Systematic patterns — cluster and quantify, do not just note. Group
cannot_testby shared signal and count each cluster: "N handlers shareStopwatch+Guid.NewGuid", "M types blocked onInternalsVisibleTo". A cluster of ≥3 is almost always ONE fix that unlocks all of them (anIClock/IGuidProvideron a shared base class, a singleInternalsVisibleTo, one extracted interface) — the highest-ROI move and the one most often missed in a long flat list. Escalate each such cluster to the human with its count and the one seam that clears it, rather than accepting N near-identical entries. Also normalize category vocabulary while here (e.g.framework_mismatch/framework-incompatible→ one label) so the report groups cleanly. Also flag any category dodged wholesale and repeated smells. - Highest-ROI next moves — concrete, not vibes.
- In-scope coverage gaps — TARGET layer only (the lens that is easiest to miss). From the
report's Risk Hotspots and per-file table, list every target-layer method/file with
any uncovered testable branch that is not in
cannot_test— uncovered branches in code we OWN. These are NOTcannot_testand NOT shallow-assertion cases (item 2 only catches covered-but-unasserted lines); they are unfinished characterization — a complex method got a happy-path test and its other branches (filters, switch arms, error paths) were never executed. A target-bucket row in Risk Hotspots is the signal. Anexcl:*hotspot is fine to leave (it is integration/E2E-covered, not unit) — only target-layer gaps count here. Close them before the baseline locks (see Output).
Trust the C0/C1 numbers; do not second-guess them — they are tool output; if a number looks wrong that is a pipeline check (file-filter, merge, instrumentation), never an LLM opinion. But trusting the numbers is NOT ignoring them: item 5 exists precisely to ACT on the in-scope gaps they reveal. (A target method at 70% slipping into the baseline because the critique was "number-blind" is the exact failure this item closes.)
Output: split findings into (a) clear-cut fixes to apply now, before baseline — a
cannot_test entry that is actually testable → write the test and remove the entry; a
target-layer coverage gap (item 5) → write the missing branch tests (run-capture-fill) until the
method's testable branches are fully covered (or the unreachable ones are in cannot_test) and it
drops off Risk Hotspots; and (b) judgment calls / larger
refactors escalated to the human. Apply the clear-cut fixes, then re-check them (loop until none
remain) — a fix can surface another, e.g. writing the test that retires one cannot_test reveals a
shared seam that retires more. The baseline does not lock while any target-layer in-scope gap
remains. Append the findings to the report under ## Suite review. Only after this loop settles
does the baseline report run.
Pick the mode
- Target existed at the baseline → characterization.
- Target is new, or a line changed after the baseline → spec.
- Mixed file → characterize baseline methods, spec the new/changed ones.
Skip anything matched by a manifest exclusions pattern — that uncovered state is
intentional and explained by the manifest.
Characterization mode — run-capture-fill (hard precondition)
For each target method:
- Write the test: explicit pinned inputs, dependencies mocked through interfaces, a placeholder expected value. Assert observable output/state/exceptions — not mock call order (interaction verification only where the side effect IS the behavior: message publish, email send, external sink with no return value).
- Run the single test. This is required. Do not write an expected value from reading the source.
- Read the actual value from the run. Write it into the assertion.
- Re-run until green.
If step 2 is genuinely impossible — needs real infrastructure with no seam, or depends on
wall-clock/id/random that flows into the asserted output with no seam — DO NOT write a
predicted assertion. Add the target (the specific method, not the file) to the manifest
cannot_test list with a category, a reason (the signal at file:line), and a mitigation
(the source change that would unlock it, or "none — genuinely nondeterministic external
boundary"). Make no source change to fix it; the backfill freezes source.
A failing-to-compile test is NOT "impossible" — it is unfinished authoring. Rule out the
authoring causes first (mock wiring, a legitimate missing test-project reference, an unstubbed
constructor argument). Only route to cannot_test when the unit cannot be constructed or
exercised through any interface after those are addressed — never on the first red build.
Cover the BRANCHES, not just one happy path — a target method is "done" only when its
branches are exercised, not when it merely has a green test. For a method with branching
(filters, a switch, error paths, asc/desc, null guards), one input lands ~one path and leaves
the rest uncovered — that is the partial-coverage gap that otherwise slips to the report's Risk
Hotspots (e.g. a GetAll with keyword + type + unit filters and a 3-way order switch needs a
test per filter and per switch arm). Enumerate the method's branches and pin an input for each;
a branch that genuinely cannot be reached without real infra goes to cannot_test, not ignored.
Do this here, at write time — it is the plan, not an afterthought. The in-loop measurement
("Measure for feedback") tells you which methods you missed; the suite critique's item 5 is only
a final backstop for what slips through, never the place you first go looking for branch coverage.
Spec mode
Expected values come from the requirement/acceptance criteria, not from running the code. The run-capture-fill loop does not apply. If a new unit is not testable through its interfaces, that is a dependency problem to fix in this change (extract interface, inject seam) — do not skip and do not widen the test project's references.
First-run triage (characterization backfill)
A characterization test is green by construction once run-capture-fill completes, because the assertion holds the actual output. The only judgment calls are the units where the captured output looks like a latent bug. Freeze them anyway (assert the actual value) and record them as observations for the report — frozen, not endorsed. Do not change behavior.
Output of a generation pass
- New/updated test files mirroring source structure, named
Method_Scenario_Expected. - Any newly discovered untestable targets appended to the manifest
cannot_testlist, each scoped to a method and carrying{ category, lines, reason, mitigation }. - A short list of suspected-latent-bug observations (target + what looked wrong), for the report's observations section. Do not fix them.
- The updated worklist checklist: N of M items done, what remains (and which phase, if phased) —
the testable set at 100% branch coverage, with the residual
cannot_testlist and its mitigation plan called out explicitly (this is the exit-gate evidence).
Only once the exit gate is met (testable set 100% branch-covered, every residual a mitigated
cannot_test entry) and the suite critique has run (see the exit/promotion gates and "Suite
critique"), run coverage-report to measure and set the baseline. Do not assert
coverage numbers yourself — they come from the tool. Until then, leave the tree dirty and stop
without committing.