Feature test
Write failing tests from work-plan contracts and acceptance criteriaFrom its SKILL.md
npx -y skills add telefrek/vallorcine --skill feature-testAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
41.1 KB, ~9.4k tokens by cl100k_base, as published. Nobody here has run it
/feature-test "<feature-slug>" [--unit <WU-N>] [--add-missing] [--escalation] [--specs <ids>]
Writes failing tests from work-plan contracts and brief acceptance criteria. Idempotent β if testing is complete for the current cycle, reports and stops. With --unit, scopes to a single work unit. With --add-missing, adds tests for cases found by the Refactor Agent. With --escalation, reviews a specific test flagged by the Code Writer as having a contract conflict. With --specs, explicit comma-separated spec IDs bypass the fuzzy-match resolver β use when the fuzzy match misses your target specs.
Idempotency pre-flight (ALWAYS FIRST)
Read .feature/<slug>/status.md.
Work unit resolution (check before anything else)
If work_units: none in status.md: ignore --unit flag, proceed as normal.
If work units are defined:
- If
--unitflag provided: scope all steps to that unit only - If no
--unitflag: find the next unit where status isnot-started(units areblockeduntil all dependencies arecomplete)- If all units complete: report and stop
- If a unit is
in-progress: resume that unit - If next unit is
blocked: display:
Stop.π§ͺ TEST WRITER Β· <slug> βββββββββββββββββββββββββββββββββββββββββββββββ WU-<n> is blocked β waiting on: <dependency units> Complete those units first, then run: /feature-test "<slug>" --unit WU-<n> Current unit statuses: <table from Work Units status>
Update the Work Units table in status.md: active unit β in-progress.
Per-unit path resolution (parallel mode)
If execution_strategy is balanced or speed in feature-level status.md:
unit_status = .feature/<slug>/units/WU-<n>/status.md
unit_log = .feature/<slug>/units/WU-<n>/cycle-log.md
Else:
unit_status = .feature/<slug>/status.md
unit_log = .feature/<slug>/cycle-log.md
All status.md reads/writes for stage, substage, cycle tracker β use unit_status.
All cycle-log.md appends β use unit_log.
Feature-level status.md Work Units table β still update unit status there too.
Display opening header with unit if applicable:
βββββββββββββββββββββββββββββββββββββββββββββββ
π§ͺ TEST WRITER Β· <slug> Β· Cycle <n>< Β· WU-<n>>
βββββββββββββββββββββββββββββββββββββββββββββββ
Determine the current TDD cycle number from the TDD Cycle Tracker.
Without --add-missing or --escalation:
First, check status.md substage for contract-revised. Also check cycle-log.md
for a recent contract-revised entry.
If a contract revision is pending: the Work Planner has revised a contract and its stub. Tests may now be misaligned with the updated contract.
- Say: "Contract revised by Work Planner β re-verifying tests against updated contract."
- Read the most recent
contract-revisedentry from cycle-log.md to understand what changed - Read the revised contract section from work-plan.md
- Read the affected test file(s)
- If any test assertions no longer match the revised contract: update the tests to align with the new contract. Run them to confirm they fail for the right reason (not-implemented, not wrong-assertion).
- If tests already match: no changes needed.
- Append
tests-reverifiedto cycle-log.md:
## <YYYY-MM-DD> β tests-reverified
**Agent:** π§ͺ Test Writer
**Cycle:** <n>
**Contract:** <construct name>
**Tests changed:** <list of changed tests, or "none β already aligned">
---
- Set status.md substage back to
complete(testing is done). - Display:
π§ͺ TEST WRITER Β· <slug> Β· contract re-verified
βββββββββββββββββββββββββββββββββββββββββββββββ
Contract: <construct name>
Tests changed: <list or "none β already aligned">
Re-run implementation:
/feature-implement "<slug>"< --unit WU-<n>>
Stop.
If Testing status for the current cycle is complete (for this unit if --unit provided):
Read automation_mode from status.md.
If automation_mode: autonomous:
π§ͺ TEST WRITER Β· <slug>
βββββββββββββββββββββββββββββββββββββββββββββββ
Tests are already written for cycle <n>< Β· WU-<n>>.
Starting implementation Β· type stop to pause
βββββββββββββββββββββββββββββββββββββββββββββββ
Invoke /feature-implement "<slug>"< --unit WU-<n>> as a sub-agent immediately.
If automation_mode: manual (or not set):
π§ͺ TEST WRITER Β· <slug>
βββββββββββββββββββββββββββββββββββββββββββββββ
Tests are already written for cycle <n>< Β· WU-<n>>.
Test plan: .feature/<slug>/test-plan.md
All tests verified failing.
In parallel mode (execution_strategy: balanced | speed): do NOT call
AskUserQuestion β chain to /feature-implement "<slug>" --unit WU-<n>
immediately. No human is attached to this subagent.
Sequential/cost mode: Use AskUserQuestion with two options:
- "Proceed to implementation"
- "Stop"
If "Proceed to implementation": invoke /feature-implement "<slug>"< --unit WU-<n>> as a sub-agent immediately.
If "Stop": display Next: /feature-implement "<slug>" and stop.
If Testing is in-progress:
- Check which test files already exist
- Say: "Test writing was in progress for cycle <n> β resuming."
- Skip tests already written; write only missing ones
- If all tests are written but failure verification hasn't run: jump to Step 4
With --add-missing:
Find the most recent missing-tests-found entry in cycle-log.md.
If none exists: "No missing tests have been reported. Nothing to add."
If found: load only the listed missing test cases and proceed directly to Step 3.
With --escalation:
Read the most recent code-escalation entry in cycle-log.md.
If none exists: "No escalation has been logged. Nothing to review." Stop.
If found: proceed to the Escalation Review section below.
If Testing is not-started:
- Verify
.feature/<slug>/work-plan.mdexists. If not: "Run /feature-plan first." - Set status.md: Testing cycle N β
in-progress, substage βplanning - Display opening header and proceed to Step 1
Display opening header:
βββββββββββββββββββββββββββββββββββββββββββββββ
π§ͺ TEST WRITER Β· <slug> Β· Cycle <n>
βββββββββββββββββββββββββββββββββββββββββββββββ
Step 0a β Progress tracking
Skip TodoWrite if execution_strategy is balanced or speed. In
parallel mode, the coordinator owns TodoWrite β subagents must not call it
or they will overwrite the coordinator's checklist.
In sequential/cost mode: use TodoWrite to show progress in the Claude Code UI (visible via Ctrl+T). Each TodoWrite call replaces the full list β always include all items.
Pipeline context: Include the full feature lifecycle as top-level items.
Mark earlier stages completed, current in_progress, later pending.
Per-test granularity: After the test plan is confirmed (Step 2), add an
item for each test case from the plan. Update each to in_progress while
writing and completed when done. Use activeForm to show which file is
being written.
Example checklist during test writing:
[
{"id": "pipeline-scoping", "content": "Scoping", "status": "completed", "priority": "medium"},
{"id": "pipeline-domains", "content": "Domain analysis", "status": "completed", "priority": "medium"},
{"id": "pipeline-planning", "content": "Work planning", "status": "completed", "priority": "medium"},
{"id": "pipeline-testing", "content": "Test writing", "status": "in_progress", "priority": "high",
"activeForm": "Writing test 3 of 8"},
{"id": "pipeline-implementation", "content": "Implementation", "status": "pending", "priority": "medium"},
{"id": "pipeline-refactor", "content": "Refactor & review", "status": "pending", "priority": "medium"},
{"id": "pipeline-pr", "content": "PR draft", "status": "pending", "priority": "medium"},
{"id": "test-1", "content": "test_returns_token_on_valid_input", "status": "completed", "priority": "high"},
{"id": "test-2", "content": "test_rejects_empty_input", "status": "completed", "priority": "high"},
{"id": "test-3", "content": "test_rate_limit_exceeded", "status": "in_progress", "priority": "high",
"activeForm": "Writing to tests/test_rate_limit.py"},
{"id": "test-4", "content": "test_concurrent_access", "status": "pending", "priority": "high"},
{"id": "verify-fail", "content": "Verify all tests fail (not-implemented)", "status": "pending", "priority": "high"},
{"id": "handoff", "content": "Hand off to implementation", "status": "pending", "priority": "medium"}
]
Step 1 β Load context
If work unit is active, load only what is needed for that unit:
.feature/project-config.mdβ always.feature/<slug>/brief.mdβ acceptance criteria and error cases.feature/<slug>/work-plan.mdβ if work units defined: read only the section for the active unit plus the Contract Definitions for its constructs. Do NOT load contract sections for other units.
Do NOT read implementation files.
If --add-missing: load only the missing test cases from cycle-log.md.
Step 1a β Resolve hardened specs (if available)
The goal is to load every relevant APPROVED spec so generated tests
carry covers: R<N> annotations tying each test to its requirement.
Three resolution paths, tried in priority order β the first one that
produces IDs wins.
1. --specs <id1,id2,β¦> flag (highest priority). If the caller
passed --specs on the command line (e.g.
/feature-test "float16-vectors" --specs query.vector-index,schema.float16-encoding),
split on commas and use those IDs directly. This bypasses fuzzy match
when the author knows exactly which specs apply.
2. .feature/<slug>/brief.md explicit list. If the brief's front
matter or narrative contains a specs: line listing IDs (e.g.
specs: [query.vector-index, schema.float16-encoding] or
specs: query.vector-index, schema.float16-encoding), extract those
IDs.
3. Fuzzy match (fallback). If neither path supplied explicit IDs, fall through to spec-resolve.sh's domain-inference matching on the feature brief text.
if [[ -n "$EXPLICIT_IDS" ]]; then
EXPLICIT_SPEC_IDS="$EXPLICIT_IDS" \
bash .claude/scripts/spec-resolve.sh "<feature brief>" 25000 2>/dev/null
else
bash .claude/scripts/spec-resolve.sh "<feature brief title or description>" 25000 2>/dev/null
fi
Store the output as SPEC_BUNDLE. When non-empty, the bundle
contains behavioral requirements (R1, R2, β¦) from hardened specs.
No silent fallback. If .spec/ exists and its manifest has at
least one APPROVED spec but the resolver returns an empty bundle,
stop immediately. Generating tests without spec annotations when
specs are available corrupts the downstream traceability β implementers
lose the covers: links they rely on. Display:
π SPEC RESOLUTION EMPTY
ββββββββββββββββββββββββββββββββββββββββββββββββ
This project uses hardened specs (.spec/ exists with APPROVED specs),
but no specs matched the feature brief "<title>".
Tests generated without spec annotations lose covers: R<N> links and
break downstream verification. Resolve this before proceeding:
1. Re-invoke with explicit IDs:
/feature-test "<slug>" --specs <id1>,<id2>
2. Or list spec IDs in .feature/<slug>/brief.md:
specs: [<id1>, <id2>]
3. Or rewrite the brief title to match the domain's vocabulary.
In parallel mode (execution_strategy: balanced | speed): do NOT call
AskUserQuestion. Append a spec-resolution-escalation entry to
units/WU-<n>/cycle-log.md (or feature-level cycle-log.md if not in a work
unit) listing the brief title and the unmatched domains, set per-unit
status.md substage β escalated-spec-resolution (or feature status.md in
non-unit mode), return:
<slug>: ESCALATED β spec-resolution: no specs matched brief "<title>"
STOP. The coordinator (or user) triages spec selection.
Sequential/cost mode: Use AskUserQuestion to offer:
- "I'll re-invoke with --specs" β stop; user re-invokes
- "Proceed without specs" β override: set SPEC_BUNDLE empty and continue (only safe if the user is certain no specs apply)
If the project has no .spec/ directory, or the manifest has no
APPROVED specs at all: set SPEC_BUNDLE empty and proceed silently β
this is a project not using the spec system, not a resolution failure.
Do not read .spec/ files directly. The resolver handles file
discovery, domain matching, transitive dependency expansion, and
token budgeting.
Step 1b β Construct analysis (derive tests from interfaces)
Read the stub files for each construct in the work plan. The stubs define
the public interface β method signatures, parameter types, return types, and
any throws/raises declarations. Derive structural test cases that the
brief and acceptance criteria don't mention but the interface implies.
For each construct, scan its stub for these patterns:
| Interface pattern | What it implies | Test to derive |
|---|---|---|
| Paired methods (encode/decode, serialize/deserialize, write/read, open/close, add/remove) | Round-trip: operation then inverse preserves data | test_<X>_then_<Y>_is_identity |
| Closeable / AutoCloseable / resource lifecycle | Cleanup on all paths, use-after-close safety | test_close_releases_resources, test_use_after_close_throws |
| Mutable state (add, put, insert, update, delete on a collection/store) | Interaction: sequences of mutations produce expected state | test_add_then_remove_restores_original, test_multiple_adds_accumulate |
| Iterator / stream / cursor | Exhaustion, empty iteration, concurrent modification | test_empty_iteration, test_iterator_exhaustion |
| Concurrency contract (from spec) | If spec declares thread-safe: concurrent access, check-then-act atomicity, close-under-contention. If spec declares not-thread-safe: document in test comment, no concurrency tests. | test_concurrent_<operation>, test_close_during_<operation> |
| Factory / builder | Invalid configuration, required fields, build order | test_build_without_required_field_throws |
| Comparable / ordering | Symmetry, transitivity, consistency with equals | test_ordering_symmetry, test_ordering_transitivity |
| Numeric parameters (capacity, size, count, limit, offset) | Zero, negative, overflow | test_zero_capacity, test_negative_throws |
| Byte buffers / arrays | Empty, single byte, exact capacity, overflow | test_empty_buffer, test_buffer_at_capacity |
| Type hierarchies (sealed types, enum switches, visitor patterns) | All variants handled | One test per variant/subtype |
How to apply: For each pattern found, add 1-2 test cases to the plan. These are structural tests β they test properties of the interface, not specific business scenarios. They go in a "Structural" section of the test plan alongside the existing Happy path, Error, and Boundary sections.
Do NOT over-generate. Only add tests for patterns actually present in the stubs. A construct with no paired methods gets no round-trip tests. The goal is to catch the 3-5 tests the refactor agent would flag, not to double the test count.
Step 1c β Spec analysis pre-pass (both lenses)
Analyze the work-plan contracts and stubs across two complementary lenses to identify risks BEFORE writing tests. This step prevents bugs from being written rather than finding them after implementation.
KB integration
If .kb/CLAUDE.md exists, scan the Topic Map for categories relevant to this
feature's domain (e.g., encryption, indexing, serialization, compression).
Read any type: adversarial-finding entries in matching categories β they
contain bug patterns discovered in prior features that may recur here. Add
matching patterns to the Lens B checklist below.
Project rules as audit vectors
Check for project-specific rules that define what "correct" means beyond
general best practices. Rules in .claude/rules/, CONTRIBUTING.md, and
accepted ADRs in .decisions/ all constrain implementation. Violations of
project rules are bugs β add them to Lens B. Examples: memory discipline
rules, architectural constraints, testing conventions, coding standards.
Lens A β Requirement operationalization / Contract gaps
Conflict pre-check: Before operationalizing requirements, check if
SPEC_BUNDLE contains a ## Conflicts section. If it does, extract all
CONFLICT and INVALIDATES lines. For each conflicting requirement pair
(e.g., F03.R8 and F07.R56), mark both requirements as:
UNTESTABLE: spec conflict β requirements <R_X> and <R_Y> contradict
Do NOT write tests for these requirements. Contradictory specs would produce contradictory tests β one asserting "must accept" while another asserts "must reject" for the same behavior. Include the UNTESTABLE entries in the test plan's defensive section so the conflict is visible but not acted upon.
Proceed with operationalizing all non-conflicting requirements as normal.
If SPEC_BUNDLE is non-empty (hardened specs available):
The bundle contains behavioral requirements (R1, R2, ...) from hardened specs. Operationalize each requirement into test cases:
- Extract every requirement ID (R1, R2, ...) and its behavioral description
from the
## Feature Requirementssection of the bundle. - For each requirement, ask: "Can I write a test that verifies this
requirement against the constructs in the work plan?"
- If yes: create one or more test cases that exercise the requirement.
Tag each test with the requirement ID it covers (e.g.,
covers: R3). - If the requirement is abstract or cross-cutting (e.g., "the system must be resilient to X"): identify which construct is responsible for that behavior and write a concrete test against it.
- If the requirement cannot be tested at the unit/integration level (e.g.,
deployment constraints): note it as
UNTESTABLE: <reason>in the defensive section. Do not silently drop it.
- If yes: create one or more test cases that exercise the requirement.
Tag each test with the requirement ID it covers (e.g.,
- After operationalizing all spec requirements, apply the contract-gap table below as a supplement β the spec may not cover every edge case the interface implies. Any gap found that isn't already covered by a spec requirement becomes a CONTRACT-GAP finding.
- Check for open obligations in the bundle's
## Open Obligationssection. Each obligation is a spec-level TODO that must be addressed. Convert applicable obligations into test cases.
If SPEC_BUNDLE is empty (no hardened specs β fallback):
For each contract in the work plan, apply the contract-gap table below as the primary analysis tool. This is the original Lens A behavior.
Contract-gap table (primary in fallback mode, supplement in spec mode):
| Question | What to test |
|---|---|
| What happens at boundary values? | Empty inputs, zero-length, max capacity, single element |
| What happens with null at every layer? | Constructor args, method params, stored fields, return values |
| Are error cases exhaustive? | Invalid combinations, inverted ranges, type mismatches |
| Are composite operations atomic? | If step A succeeds but step B fails, what state? |
| Are mutable inputs defensively copied? | Arrays, collections crossing trust boundaries |
| What equality semantics do keys use? | Identity vs content (especially byte[], arrays) |
Lens B β Implementation risk patterns (what code typically gets wrong)
Spec-aware scoping: If SPEC_BUNDLE is non-empty, use the spec requirements as the boundary of what is "specified behavior." Lens B then focuses on:
- Behaviors the spec did not anticipate β interactions, edge cases, and failure modes that fall outside any R_N requirement
- Gaps between requirements β where two requirements interact but neither fully specifies the combined behavior
- Implementation assumptions the spec takes for granted (e.g., thread safety, ordering, resource cleanup) without an explicit requirement
When specs are available, tag each Lens B finding with whether it is:
SPEC-BOUNDARYβ the spec has a relevant requirement but doesn't cover this specific edge caseSPEC-BLIND-SPOTβ no spec requirement addresses this area at allIMPL-RISKβ standard implementation risk (same as no-spec mode)
If SPEC_BUNDLE is empty, proceed with the standard implementation risk analysis below without spec-boundary tagging.
For each construct in scope, trace the full data flow β not just the construct itself but its inputs, outputs, and data carriers. This prevents multi-pass discovery where each audit finds the next layer.
Level 1 β The construct itself:
byte[]or arrays used as map/set keys β identity equality, not content- Mutable arrays/collections stored by reference without defensive copying
- Float/double encoding β sign-bit handling differences between integer and IEEE 754
- Multi-step mutations that aren't atomic β delete-then-insert, check-then-act
- Concurrency contract violations β if the spec declares this construct thread-safe, check: are compound operations (get-or-create, check-then-act, position-then-read) actually atomic? If the spec declares not-thread-safe, note it but don't generate false concurrency findings
- Switch/instanceof that don't cover all sealed interface subtypes
- Silent truncation or Math.min instead of fail-fast on mismatched dimensions/sizes
- Not-equals predicates interacting with null field values
- Resource lifecycle β double-close, use-after-close, deferred exception aggregation
- Validation that should happen at construction but is deferred to usage
- Any patterns from adversarial KB entries loaded above
Level 2 β Inputs (who calls this construct, what do they pass?):
- Are callers validated at the trust boundary, or is invalid input silently accepted?
- Per project rules, should out-of-range values be rejected at entry rather than handled downstream? (fail-fast principle)
- Can callers pass values that are technically valid but semantically wrong? (e.g., NaN as a score, negative capacity, inverted range bounds)
Level 3 β Outputs (what does this construct return, can consumers misuse it?):
- Do returned references expose mutable internal state? (check accessors, not just constructors)
- Are returned collections unmodifiable, or can consumers corrupt internal state?
- Can the return value be in a state the consumer doesn't expect? (null, empty, partial)
Level 4 β Data carriers (records, DTOs, result types):
- Do data carrier types enforce their own invariants at construction? (null fields, NaN scores, negative counts, empty required fields)
- Do records with mutable fields (arrays, collections, MemorySegment) have correct equals/hashCode? (identity vs content semantics)
- Are carriers immutable once constructed, or can state leak through accessors?
Scoping: Trace all 4 levels on every construct in the current work unit. If the work unit has many constructs, prioritize depth on constructs flagged by Lens A or Level 1 findings, but do not skip levels entirely.
Output
For each finding, note:
- The construct and contract section it applies to
- The finding type:
SPEC-REQβ directly operationalized from a spec requirement (Lens A, spec mode)CONTRACT-GAPβ gap not covered by any spec requirement (Lens A)SPEC-BOUNDARYβ spec-adjacent edge case (Lens B, spec mode)SPEC-BLIND-SPOTβ no spec coverage at all (Lens B, spec mode)IMPL-RISKβ standard implementation risk (Lens B)DISPLACEMENTβ behavior that must be verified as removed (from displacement resolution)
- A specific defensive test case to add to the test plan
- If from a spec requirement: the requirement ID (e.g.,
R3) - If from displacement: the displaced spec and requirement ID (e.g.,
displaced: F05.R3)
These findings feed directly into the test plan as a "Defensive (from spec analysis)" section. Do NOT write a separate spec-analysis.md file β the findings are integrated into the test plan in the next step.
Displacement verification
If the work-plan contains a ## Removal Work section, derive negative tests
for each removal work unit. These verify that displaced behavior is gone:
For each removal entry (RW-1, RW-2, ...):
- Read the displaced spec requirement text to understand what behavior existed
- Derive a test that asserts the behavior no longer exists:
- API removed β calling it throws/returns an error
- Format no longer supported β input in old format is rejected with a clear error
- Configuration removed β using the old config produces a validation error
- Behavioral change β old behavior demonstrably not happening
- Tag as
covers: <existing_id>.<req_id> (displaced)and finding typeDISPLACEMENT
These tests go in the "Displacement verification" section of the test plan (between "Spec requirements" and "Defensive").
Step 2 β Write the test plan (in chat first)
Update status.md substage β confirming-plan.
Coverage checklist (internal β apply before presenting the plan)
Before presenting the test plan, systematically check each construct against these categories. The refactor agent (step 2e) will check these exact categories later β gaps found there trigger an escalation cycle. Catching them here is much cheaper.
For each construct in the work plan:
| Category | What to test | Common gaps |
|---|---|---|
| Happy path | Normal operation with valid inputs | Rarely missed |
| Error conditions | Every error case in the brief + contract | Missing: errors not in brief but implied by types |
| Boundary values | Empty/zero, single element, max capacity, null/nil | Most commonly missed β add at least one per construct |
| State transitions | Invalid state (e.g., use after close, double init) | Missed when construct has lifecycle |
| Concurrency | Thread safety, ordering dependencies, race conditions | Only if construct is shared/concurrent β skip if single-threaded |
| Type boundaries | Overflow, underflow, precision loss, encoding limits | Missed for numeric types, byte buffers, serialization |
The most commonly missed category is boundary values. For every construct, ask: "what happens at empty, at one, and at max?" If the answer isn't obvious from the contract, write a test for it.
Present the plan
Display:
ββ Test plan βββββββββββββββββββββββββββββββββββ
TEST PLAN β <slug> (Cycle <n>)
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Happy path
1. test_<name> β <scenario> β covers: <acceptance criterion>
...
Error and edge cases
N. test_<name> β <scenario>
...
Boundary values
N. test_<name> β <scenario>
...
Structural (from interface analysis)
N. test_<name> β <scenario> β pattern: <round-trip | lifecycle | interaction | ...>
...
Spec requirements (from hardened specs) β only when specs loaded
N. test_<name> β <scenario> β covers: <spec-id>.R<N>
...
Displacement verification (removal tests) β only when removal work exists
N. test_<name> β <scenario> β covers: <spec-id>.R<N> (displaced)
...
Defensive (from spec analysis)
N. test_<name> β <scenario> β finding: <CONTRACT-GAP | SPEC-BOUNDARY | SPEC-BLIND-SPOT | IMPL-RISK>: <description>
...
βββββββββββββββββββββββββββββββββββββββββββββββ
Spec analysis: <N> SPEC-REQ + <N> CONTRACT-GAP + <N> IMPL-RISK findings β <N> defensive tests
[if specs loaded: <N> requirements operationalized, <N> untestable]
Does this cover the acceptance criteria? Any to add or remove?
βββββββββββββββββββββββββββββββββββββββββββββββ
The plan MUST include "Boundary values", "Structural", and "Defensive" sections. If any section has no applicable tests (rare), state why explicitly. The structural section should list which interface patterns were checked and found not applicable. The defensive section should summarize how many findings came from each lens and any adversarial KB patterns that were checked.
When specs were loaded: The defensive section should additionally report:
- How many spec requirements were operationalized (SPEC-REQ count)
- How many requirements were marked UNTESTABLE and why
- How many SPEC-BOUNDARY and SPEC-BLIND-SPOT findings were identified
- Which requirement IDs each defensive test covers
Wait for confirmation. Update status.md substage β writing-tests after confirm.
Step 3 β Write test files
Write to the test directory from project-config.md.
Idempotent: check whether each test already exists before writing. Append to existing test files rather than overwriting; do not duplicate tests. Before editing an existing test file, re-read it to pick up any additions from prior test-writing passes β stale reads cause Edit old_string mismatches or silent overwrites. After writing each test method, re-read the file to verify the method is present.
Rules:
- Test names describe behaviour:
test_returns_error_when_input_is_empty - Public interface only β no reaching into implementation details
- Mocks/fakes for all external dependencies
- Every test: arrange / act / assert clearly separated
- Every construct MUST have at least one boundary value test (empty, zero, max, null/nil, single element) β the refactor agent checks for these and will escalate if missing
- Every test MUST carry an
@specannotation pointing at the requirement it satisfies β seerules/spec-annotation-protocol.mdfor the format. Use the test plan'scovers: <spec-id>.R<n>line as the write-time source: translate it directly into the test file's comment syntax above the test method. Skip only when no spec was loaded for this feature (the coverage table will be vacuous in that case).
Update status.md substage β verifying-failures after writing.
Step 3a β Update spec coverage (mandatory if specs were loaded)
Run the coverage updater so the Test column of
.feature/<slug>/spec-coverage.md reflects the annotations you just
wrote. This is the structural enforcement of the annotation protocol β
without it, /feature-pr cannot tell which requirements are bound.
if [[ -f .feature/<slug>/spec-coverage.md ]]; then
bash .claude/scripts/spec-coverage.sh update \
.feature/<slug>/spec-coverage.md --all
fi
Exit precondition. Read back the coverage file. Every row whose
covers: entry in the test plan named a requirement you just wrote a
test for MUST now show a non-pending value in the Test column. If
any does not:
- Open the test file and confirm the
@specline is present, formatted correctly, and references the right<spec-id>.R<n>. - Re-run the updater.
- If the cell still says
pending, the annotation is malformed βspec-trace.shdid not match it. Fix the annotation and repeat.
Do not advance to Step 4 until every test row that should be annotated is annotated. The pipeline relies on this β partial annotation here becomes a silent gate failure at PR time.
Step 4 β Verify tests fail
Run the test suite (5-minute Bash timeout per tdd-protocol). If the suite times out: run individual test methods to isolate which test is hanging. For hanging tests, add a @Timeout annotation or rewrite with a non-blocking approach. Do not retry the full suite without isolating first.
Expected: all new tests fail with NotImplementedError or import/compile error.
If a test PASSES unexpectedly: investigate β stub may already be implemented, or test may not be testing the right thing. Do not hand off with passing tests at this stage.
Capture the full test runner output.
Escalation Review (--escalation only)
Entered when the Code Writer escalates a contract conflict via --escalation.
Step E1 β Load the escalation
Read the most recent code-escalation entry from cycle-log.md. Extract:
- The test name and file
- What the test expects
- The constraint from the work plan
- The conflict description
- The escalation count (N of 3)
Read the test file and the relevant contract section from work-plan.md.
Step E1a β Check for spec conflict
If the escalation entry's conflict description contains "SPEC CONFLICT" or
the substage in status.md is spec-conflict-detected, this is a requirement
contradiction β not a test or contract problem.
Do NOT rewrite the test. Instead:
-
Mark both the passing and failing tests as BLOCKED in cycle-log.md:
## <YYYY-MM-DD> β tests-blocked-spec-conflict **Agent:** π§ͺ Test Writer **Cycle:** <n> **Blocked tests:** `<passing test>`, `<failing test>` **Reason:** Contradictory spec requirements β <R_N> vs <R_N> **Resolution:** Requires /spec-author to reconcile conflicting requirements --- -
Display:
π§ͺ TEST WRITER Β· spec conflict Β· <slug> βββββββββββββββββββββββββββββββββββββββββββββββ This escalation is a spec conflict, not a test or contract problem. Both tests are correct given their respective requirements β the requirements themselves contradict each other. Blocked tests: - <passing test> (covers: <R_N>) - <failing test> (covers: <R_N>) Run /spec-author to resolve the conflicting requirements, then re-run /feature-test "<slug>" to unblock. -
Stop. Do not proceed to Step E2.
Step E2 β Diagnose
Determine which of three cases applies:
-
Test is wrong β the test asserts something not implied by the contract. Fix the test to match the contract. Run it to confirm it fails for the right reason (not-implemented, not wrong-assertion).
-
Test is right, contract is ambiguous β the contract can be read multiple ways. Clarify the test (add a comment explaining intent) and adjust the assertion if needed. The Code Writer will re-read the test on next run.
-
Contract itself is wrong β the work plan constraint contradicts the brief or an ADR, or is internally inconsistent. The Test Writer cannot fix this. Escalate to the Work Planner β see Step E3.
For cases 1 and 2, after fixing:
Append test-escalation-resolved to cycle-log.md:
## <YYYY-MM-DD> β test-escalation-resolved
**Agent:** π§ͺ Test Writer
**Cycle:** <n>
**Test:** `<test name>` in `<file>`
**Diagnosis:** <test was wrong | contract was ambiguous>
**Fix:** <what changed>
**Escalation count:** <N> of 3
---
Update status.md substage β escalation-resolved.
Display:
π§ͺ TEST WRITER Β· escalation resolved Β· <slug>
βββββββββββββββββββββββββββββββββββββββββββββββ
Test: <test name>
Diagnosis: <test was wrong | contract was ambiguous>
Fix: <what changed>
Re-run implementation:
/feature-implement "<slug>"< --unit WU-<n>>
Stop.
Step E3 β Escalate to Work Planner
Before escalating, check the escalation counter.
Read cycle-log.md and count test-to-planner-escalation entries for the same
contract/construct.
3rd escalation on the same contract: hard stop. Do NOT escalate to the Work Planner again. Instead:
π ESCALATION LIMIT Β· Test Writer β Manual Resolution
βββββββββββββββββββββββββββββββββββββββββββββββ
The same contract issue has been escalated 3 times without resolution.
Contract: <construct name>
Work plan section: <reference>
Automatic resolution is not working. Please review the conflict manually:
1. Check the contract in work-plan.md
2. Check the acceptance criteria in brief.md
3. Check any governing ADRs
4. Fix the work plan, then re-run /feature-test "<slug>"
If the brief itself is wrong, revisit /feature "<slug>".
Update status.md substage β escalation-limit-reached. Stop.
Under the limit: proceed with escalation.
Append test-to-planner-escalation to cycle-log.md:
## <YYYY-MM-DD> β test-to-planner-escalation
**Agent:** π§ͺ Test Writer
**Cycle:** <n>
**Contract:** <construct name>
**Work plan section:** <reference>
**Conflict:** <what the contract says vs. what it should say>
**Brief reference:** <acceptance criterion that contradicts the contract>
**Escalation count:** <N> of 3
---
Update status.md substage β escalated-to-work-planner.
Display:
β οΈ ESCALATION Β· Test Writer β Work Planner (<N>/3)
βββββββββββββββββββββββββββββββββββββββββββββββ
Contract conflict β the work plan constraint cannot satisfy the brief.
Contract: <construct name>
Problem: <paragraph>
Brief reference: <acceptance criterion>
The Work Planner needs to revise this contract. Run:
/feature-plan "<slug>"
Then re-run the test β implement cycle for this construct.
Stop.
Step 5 β Write test-plan.md and log
Write .feature/<slug>/test-plan.md (or append cycle section if it exists):
## Cycle <n> β <YYYY-MM-DD>
### Tests Written
| Test name | File | Construct | Acceptance criterion |
|-----------|------|-----------|---------------------|
| test_<n> | <path> | <construct> | <criterion> |
### Failure output (expected)
<test runner output>
```
Coverage
- Acceptance criteria covered: <n>/<total>
- Error cases covered: <n>
- Spec requirements operationalized: <n>/<total> (omit if no specs loaded)
- Untestable requirements: <list or none> (omit if no specs loaded)
- Gaps noted: <any>
Update status.md:
- Testing cycle N β `complete`
- TDD Cycle Tracker: Tests written β today
- substage β "tests verified failing"
- Stage Completion table: Testing row β Est. Tokens `~<N>K` (project-config ~1K +
brief ~2K + work-plan section ~2K + test files written)
Append `tests-written` entry to cycle-log.md.
Update `.feature/CLAUDE.md`.
---
## Step 6 β Hand off
Read `automation_mode` from status.md.
### Determine next stage: hardening or implement
Check the work-plan contracts for domain lens signals to decide whether
hardening is needed. Count constructs with any of these contract properties:
- Closeable/AutoCloseable, close/cleanup mentions β resource_lifecycle
- Cross-module dependencies β contract_boundaries
- Thread-safety mentions, `shares_state` edges β concurrency
- Encode/decode, serialize/deserialize β data_transformation
- Mutable state shared by 2+ constructs β shared_state
**If zero lens signals OR only 1 construct in the work plan:**
- Skip hardening β chain directly to `/feature-implement`.
**If 1-2 lens signals AND 2-5 constructs:**
- Chain to `/feature-harden "<slug>" --lite< --unit WU-<n>>`.
**If 3+ lens signals OR 6+ constructs:**
- Chain to `/feature-harden "<slug>"< --unit WU-<n>>`.
### Chain execution
**If `automation_mode: autonomous`:**
Display the summary then chain immediately without prompting:
βββββββββββββββββββββββββββββββββββββββββββββββ π§ͺ TEST WRITER complete Β· <slug> Β· Cycle <n>< Β· WU-<n>> Tokens : <TOKEN_USAGE> βββββββββββββββββββββββββββββββββββββββββββββββ Tests written and verified failing. Cycle <n>< Β· WU-<n>>.
Starting <hardening | implementation> Β· type stop to pause βββββββββββββββββββββββββββββββββββββββββββββββ
Then invoke the next stage as a sub-agent immediately.
**If `automation_mode: manual` (or not set):**
Display:
βββββββββββββββββββββββββββββββββββββββββββββββ π§ͺ TEST WRITER complete Β· <slug> Β· Cycle <n>< Β· WU-<n>> Tokens : <TOKEN_USAGE> βββββββββββββββββββββββββββββββββββββββββββββββ Tests written and verified failing. Cycle <n>< Β· WU-<n>>.
Use AskUserQuestion with two options:
- "Continue"
- "Stop"
If "Continue": invoke the next stage as a sub-agent immediately.
If "Stop":
When you're ready: /feature-harden "<slug>"< --unit WU-<n>> (or /feature-implement if skipping)
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.