agentsclimarketplace

Feature test

Skill telefrek/vallorcine/skills/feature-test

Write failing tests from work-plan contracts and acceptance criteriaFrom its SKILL.md

Install
npx -y skills add telefrek/vallorcine --skill feature-test

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

41.1 KB, ~9.4k tokens by cl100k_base, as published. Nobody here has run it

/feature-test "<feature-slug>" [--unit <WU-N>] [--add-missing] [--escalation] [--specs <ids>]

Writes failing tests from work-plan contracts and brief acceptance criteria. Idempotent β€” if testing is complete for the current cycle, reports and stops. With --unit, scopes to a single work unit. With --add-missing, adds tests for cases found by the Refactor Agent. With --escalation, reviews a specific test flagged by the Code Writer as having a contract conflict. With --specs, explicit comma-separated spec IDs bypass the fuzzy-match resolver β€” use when the fuzzy match misses your target specs.


Idempotency pre-flight (ALWAYS FIRST)

Read .feature/<slug>/status.md.

Work unit resolution (check before anything else)

If work_units: none in status.md: ignore --unit flag, proceed as normal.

If work units are defined:

  • If --unit flag provided: scope all steps to that unit only
  • If no --unit flag: find the next unit where status is not-started (units are blocked until all dependencies are complete)
    • If all units complete: report and stop
    • If a unit is in-progress: resume that unit
    • If next unit is blocked: display:
      πŸ§ͺ TEST WRITER Β· <slug>
      ───────────────────────────────────────────────
      WU-<n> is blocked β€” waiting on: <dependency units>
      Complete those units first, then run:
        /feature-test "<slug>" --unit WU-<n>
      Current unit statuses:
        <table from Work Units status>
      
      Stop.

Update the Work Units table in status.md: active unit β†’ in-progress.

Per-unit path resolution (parallel mode)

If execution_strategy is balanced or speed in feature-level status.md:

unit_status = .feature/<slug>/units/WU-<n>/status.md
unit_log = .feature/<slug>/units/WU-<n>/cycle-log.md

Else:

unit_status = .feature/<slug>/status.md
unit_log = .feature/<slug>/cycle-log.md

All status.md reads/writes for stage, substage, cycle tracker β†’ use unit_status. All cycle-log.md appends β†’ use unit_log. Feature-level status.md Work Units table β†’ still update unit status there too.

Display opening header with unit if applicable:

───────────────────────────────────────────────
πŸ§ͺ TEST WRITER Β· <slug> Β· Cycle <n><  Β· WU-<n>>
───────────────────────────────────────────────

Determine the current TDD cycle number from the TDD Cycle Tracker.

Without --add-missing or --escalation:

First, check status.md substage for contract-revised. Also check cycle-log.md for a recent contract-revised entry.

If a contract revision is pending: the Work Planner has revised a contract and its stub. Tests may now be misaligned with the updated contract.

  • Say: "Contract revised by Work Planner β€” re-verifying tests against updated contract."
  • Read the most recent contract-revised entry from cycle-log.md to understand what changed
  • Read the revised contract section from work-plan.md
  • Read the affected test file(s)
  • If any test assertions no longer match the revised contract: update the tests to align with the new contract. Run them to confirm they fail for the right reason (not-implemented, not wrong-assertion).
  • If tests already match: no changes needed.
  • Append tests-reverified to cycle-log.md:
## <YYYY-MM-DD> β€” tests-reverified
**Agent:** πŸ§ͺ Test Writer
**Cycle:** <n>
**Contract:** <construct name>
**Tests changed:** <list of changed tests, or "none β€” already aligned">
---
  • Set status.md substage back to complete (testing is done).
  • Display:
πŸ§ͺ TEST WRITER Β· <slug> Β· contract re-verified
───────────────────────────────────────────────
Contract: <construct name>
Tests changed: <list or "none β€” already aligned">

Re-run implementation:
  /feature-implement "<slug>"<  --unit WU-<n>>

Stop.

If Testing status for the current cycle is complete (for this unit if --unit provided):

Read automation_mode from status.md.

If automation_mode: autonomous:

πŸ§ͺ TEST WRITER Β· <slug>
───────────────────────────────────────────────
Tests are already written for cycle <n><  Β· WU-<n>>.
Starting implementation  Β·  type stop to pause
───────────────────────────────────────────────

Invoke /feature-implement "<slug>"< --unit WU-<n>> as a sub-agent immediately.

If automation_mode: manual (or not set):

πŸ§ͺ TEST WRITER Β· <slug>
───────────────────────────────────────────────
Tests are already written for cycle <n><  Β· WU-<n>>.
Test plan: .feature/<slug>/test-plan.md
All tests verified failing.

In parallel mode (execution_strategy: balanced | speed): do NOT call AskUserQuestion β€” chain to /feature-implement "<slug>" --unit WU-<n> immediately. No human is attached to this subagent.

Sequential/cost mode: Use AskUserQuestion with two options:

  • "Proceed to implementation"
  • "Stop"

If "Proceed to implementation": invoke /feature-implement "<slug>"< --unit WU-<n>> as a sub-agent immediately. If "Stop": display Next: /feature-implement "<slug>" and stop.

If Testing is in-progress:

  • Check which test files already exist
  • Say: "Test writing was in progress for cycle <n> β€” resuming."
  • Skip tests already written; write only missing ones
  • If all tests are written but failure verification hasn't run: jump to Step 4

With --add-missing:

Find the most recent missing-tests-found entry in cycle-log.md. If none exists: "No missing tests have been reported. Nothing to add." If found: load only the listed missing test cases and proceed directly to Step 3.

With --escalation:

Read the most recent code-escalation entry in cycle-log.md. If none exists: "No escalation has been logged. Nothing to review." Stop. If found: proceed to the Escalation Review section below.

If Testing is not-started:

  • Verify .feature/<slug>/work-plan.md exists. If not: "Run /feature-plan first."
  • Set status.md: Testing cycle N β†’ in-progress, substage β†’ planning
  • Display opening header and proceed to Step 1

Display opening header:

───────────────────────────────────────────────
πŸ§ͺ TEST WRITER Β· <slug> Β· Cycle <n>
───────────────────────────────────────────────

Step 0a β€” Progress tracking

Skip TodoWrite if execution_strategy is balanced or speed. In parallel mode, the coordinator owns TodoWrite β€” subagents must not call it or they will overwrite the coordinator's checklist.

In sequential/cost mode: use TodoWrite to show progress in the Claude Code UI (visible via Ctrl+T). Each TodoWrite call replaces the full list β€” always include all items.

Pipeline context: Include the full feature lifecycle as top-level items. Mark earlier stages completed, current in_progress, later pending.

Per-test granularity: After the test plan is confirmed (Step 2), add an item for each test case from the plan. Update each to in_progress while writing and completed when done. Use activeForm to show which file is being written.

Example checklist during test writing:

[
  {"id": "pipeline-scoping", "content": "Scoping", "status": "completed", "priority": "medium"},
  {"id": "pipeline-domains", "content": "Domain analysis", "status": "completed", "priority": "medium"},
  {"id": "pipeline-planning", "content": "Work planning", "status": "completed", "priority": "medium"},
  {"id": "pipeline-testing", "content": "Test writing", "status": "in_progress", "priority": "high",
   "activeForm": "Writing test 3 of 8"},
  {"id": "pipeline-implementation", "content": "Implementation", "status": "pending", "priority": "medium"},
  {"id": "pipeline-refactor", "content": "Refactor & review", "status": "pending", "priority": "medium"},
  {"id": "pipeline-pr", "content": "PR draft", "status": "pending", "priority": "medium"},
  {"id": "test-1", "content": "test_returns_token_on_valid_input", "status": "completed", "priority": "high"},
  {"id": "test-2", "content": "test_rejects_empty_input", "status": "completed", "priority": "high"},
  {"id": "test-3", "content": "test_rate_limit_exceeded", "status": "in_progress", "priority": "high",
   "activeForm": "Writing to tests/test_rate_limit.py"},
  {"id": "test-4", "content": "test_concurrent_access", "status": "pending", "priority": "high"},
  {"id": "verify-fail", "content": "Verify all tests fail (not-implemented)", "status": "pending", "priority": "high"},
  {"id": "handoff", "content": "Hand off to implementation", "status": "pending", "priority": "medium"}
]


Step 1 β€” Load context

If work unit is active, load only what is needed for that unit:

  1. .feature/project-config.md β€” always
  2. .feature/<slug>/brief.md β€” acceptance criteria and error cases
  3. .feature/<slug>/work-plan.md β€” if work units defined: read only the section for the active unit plus the Contract Definitions for its constructs. Do NOT load contract sections for other units.

Do NOT read implementation files.

If --add-missing: load only the missing test cases from cycle-log.md.

Step 1a β€” Resolve hardened specs (if available)

The goal is to load every relevant APPROVED spec so generated tests carry covers: R<N> annotations tying each test to its requirement. Three resolution paths, tried in priority order β€” the first one that produces IDs wins.

1. --specs <id1,id2,…> flag (highest priority). If the caller passed --specs on the command line (e.g. /feature-test "float16-vectors" --specs query.vector-index,schema.float16-encoding), split on commas and use those IDs directly. This bypasses fuzzy match when the author knows exactly which specs apply.

2. .feature/<slug>/brief.md explicit list. If the brief's front matter or narrative contains a specs: line listing IDs (e.g. specs: [query.vector-index, schema.float16-encoding] or specs: query.vector-index, schema.float16-encoding), extract those IDs.

3. Fuzzy match (fallback). If neither path supplied explicit IDs, fall through to spec-resolve.sh's domain-inference matching on the feature brief text.

if [[ -n "$EXPLICIT_IDS" ]]; then
  EXPLICIT_SPEC_IDS="$EXPLICIT_IDS" \
    bash .claude/scripts/spec-resolve.sh "<feature brief>" 25000 2>/dev/null
else
  bash .claude/scripts/spec-resolve.sh "<feature brief title or description>" 25000 2>/dev/null
fi

Store the output as SPEC_BUNDLE. When non-empty, the bundle contains behavioral requirements (R1, R2, …) from hardened specs.

No silent fallback. If .spec/ exists and its manifest has at least one APPROVED spec but the resolver returns an empty bundle, stop immediately. Generating tests without spec annotations when specs are available corrupts the downstream traceability β€” implementers lose the covers: links they rely on. Display:

πŸ›‘  SPEC RESOLUTION EMPTY
────────────────────────────────────────────────
This project uses hardened specs (.spec/ exists with APPROVED specs),
but no specs matched the feature brief "<title>".

Tests generated without spec annotations lose covers: R<N> links and
break downstream verification. Resolve this before proceeding:

  1. Re-invoke with explicit IDs:
     /feature-test "<slug>" --specs <id1>,<id2>

  2. Or list spec IDs in .feature/<slug>/brief.md:
     specs: [<id1>, <id2>]

  3. Or rewrite the brief title to match the domain's vocabulary.

In parallel mode (execution_strategy: balanced | speed): do NOT call AskUserQuestion. Append a spec-resolution-escalation entry to units/WU-<n>/cycle-log.md (or feature-level cycle-log.md if not in a work unit) listing the brief title and the unmatched domains, set per-unit status.md substage β†’ escalated-spec-resolution (or feature status.md in non-unit mode), return:

<slug>: ESCALATED β€” spec-resolution: no specs matched brief "<title>"

STOP. The coordinator (or user) triages spec selection.

Sequential/cost mode: Use AskUserQuestion to offer:

  • "I'll re-invoke with --specs" β€” stop; user re-invokes
  • "Proceed without specs" β€” override: set SPEC_BUNDLE empty and continue (only safe if the user is certain no specs apply)

If the project has no .spec/ directory, or the manifest has no APPROVED specs at all: set SPEC_BUNDLE empty and proceed silently β€” this is a project not using the spec system, not a resolution failure.

Do not read .spec/ files directly. The resolver handles file discovery, domain matching, transitive dependency expansion, and token budgeting.


Step 1b β€” Construct analysis (derive tests from interfaces)

Read the stub files for each construct in the work plan. The stubs define the public interface β€” method signatures, parameter types, return types, and any throws/raises declarations. Derive structural test cases that the brief and acceptance criteria don't mention but the interface implies.

For each construct, scan its stub for these patterns:

Interface patternWhat it impliesTest to derive
Paired methods (encode/decode, serialize/deserialize, write/read, open/close, add/remove)Round-trip: operation then inverse preserves datatest_<X>_then_<Y>_is_identity
Closeable / AutoCloseable / resource lifecycleCleanup on all paths, use-after-close safetytest_close_releases_resources, test_use_after_close_throws
Mutable state (add, put, insert, update, delete on a collection/store)Interaction: sequences of mutations produce expected statetest_add_then_remove_restores_original, test_multiple_adds_accumulate
Iterator / stream / cursorExhaustion, empty iteration, concurrent modificationtest_empty_iteration, test_iterator_exhaustion
Concurrency contract (from spec)If spec declares thread-safe: concurrent access, check-then-act atomicity, close-under-contention. If spec declares not-thread-safe: document in test comment, no concurrency tests.test_concurrent_<operation>, test_close_during_<operation>
Factory / builderInvalid configuration, required fields, build ordertest_build_without_required_field_throws
Comparable / orderingSymmetry, transitivity, consistency with equalstest_ordering_symmetry, test_ordering_transitivity
Numeric parameters (capacity, size, count, limit, offset)Zero, negative, overflowtest_zero_capacity, test_negative_throws
Byte buffers / arraysEmpty, single byte, exact capacity, overflowtest_empty_buffer, test_buffer_at_capacity
Type hierarchies (sealed types, enum switches, visitor patterns)All variants handledOne test per variant/subtype

How to apply: For each pattern found, add 1-2 test cases to the plan. These are structural tests β€” they test properties of the interface, not specific business scenarios. They go in a "Structural" section of the test plan alongside the existing Happy path, Error, and Boundary sections.

Do NOT over-generate. Only add tests for patterns actually present in the stubs. A construct with no paired methods gets no round-trip tests. The goal is to catch the 3-5 tests the refactor agent would flag, not to double the test count.


Step 1c β€” Spec analysis pre-pass (both lenses)

Analyze the work-plan contracts and stubs across two complementary lenses to identify risks BEFORE writing tests. This step prevents bugs from being written rather than finding them after implementation.

KB integration

If .kb/CLAUDE.md exists, scan the Topic Map for categories relevant to this feature's domain (e.g., encryption, indexing, serialization, compression). Read any type: adversarial-finding entries in matching categories β€” they contain bug patterns discovered in prior features that may recur here. Add matching patterns to the Lens B checklist below.

Project rules as audit vectors

Check for project-specific rules that define what "correct" means beyond general best practices. Rules in .claude/rules/, CONTRIBUTING.md, and accepted ADRs in .decisions/ all constrain implementation. Violations of project rules are bugs β€” add them to Lens B. Examples: memory discipline rules, architectural constraints, testing conventions, coding standards.

Lens A β€” Requirement operationalization / Contract gaps

Conflict pre-check: Before operationalizing requirements, check if SPEC_BUNDLE contains a ## Conflicts section. If it does, extract all CONFLICT and INVALIDATES lines. For each conflicting requirement pair (e.g., F03.R8 and F07.R56), mark both requirements as:

UNTESTABLE: spec conflict β€” requirements <R_X> and <R_Y> contradict

Do NOT write tests for these requirements. Contradictory specs would produce contradictory tests β€” one asserting "must accept" while another asserts "must reject" for the same behavior. Include the UNTESTABLE entries in the test plan's defensive section so the conflict is visible but not acted upon.

Proceed with operationalizing all non-conflicting requirements as normal.

If SPEC_BUNDLE is non-empty (hardened specs available):

The bundle contains behavioral requirements (R1, R2, ...) from hardened specs. Operationalize each requirement into test cases:

  1. Extract every requirement ID (R1, R2, ...) and its behavioral description from the ## Feature Requirements section of the bundle.
  2. For each requirement, ask: "Can I write a test that verifies this requirement against the constructs in the work plan?"
    • If yes: create one or more test cases that exercise the requirement. Tag each test with the requirement ID it covers (e.g., covers: R3).
    • If the requirement is abstract or cross-cutting (e.g., "the system must be resilient to X"): identify which construct is responsible for that behavior and write a concrete test against it.
    • If the requirement cannot be tested at the unit/integration level (e.g., deployment constraints): note it as UNTESTABLE: <reason> in the defensive section. Do not silently drop it.
  3. After operationalizing all spec requirements, apply the contract-gap table below as a supplement β€” the spec may not cover every edge case the interface implies. Any gap found that isn't already covered by a spec requirement becomes a CONTRACT-GAP finding.
  4. Check for open obligations in the bundle's ## Open Obligations section. Each obligation is a spec-level TODO that must be addressed. Convert applicable obligations into test cases.

If SPEC_BUNDLE is empty (no hardened specs β€” fallback):

For each contract in the work plan, apply the contract-gap table below as the primary analysis tool. This is the original Lens A behavior.

Contract-gap table (primary in fallback mode, supplement in spec mode):

QuestionWhat to test
What happens at boundary values?Empty inputs, zero-length, max capacity, single element
What happens with null at every layer?Constructor args, method params, stored fields, return values
Are error cases exhaustive?Invalid combinations, inverted ranges, type mismatches
Are composite operations atomic?If step A succeeds but step B fails, what state?
Are mutable inputs defensively copied?Arrays, collections crossing trust boundaries
What equality semantics do keys use?Identity vs content (especially byte[], arrays)

Lens B β€” Implementation risk patterns (what code typically gets wrong)

Spec-aware scoping: If SPEC_BUNDLE is non-empty, use the spec requirements as the boundary of what is "specified behavior." Lens B then focuses on:

  • Behaviors the spec did not anticipate β€” interactions, edge cases, and failure modes that fall outside any R_N requirement
  • Gaps between requirements β€” where two requirements interact but neither fully specifies the combined behavior
  • Implementation assumptions the spec takes for granted (e.g., thread safety, ordering, resource cleanup) without an explicit requirement

When specs are available, tag each Lens B finding with whether it is:

  • SPEC-BOUNDARY β€” the spec has a relevant requirement but doesn't cover this specific edge case
  • SPEC-BLIND-SPOT β€” no spec requirement addresses this area at all
  • IMPL-RISK β€” standard implementation risk (same as no-spec mode)

If SPEC_BUNDLE is empty, proceed with the standard implementation risk analysis below without spec-boundary tagging.

For each construct in scope, trace the full data flow β€” not just the construct itself but its inputs, outputs, and data carriers. This prevents multi-pass discovery where each audit finds the next layer.

Level 1 β€” The construct itself:

  • byte[] or arrays used as map/set keys β€” identity equality, not content
  • Mutable arrays/collections stored by reference without defensive copying
  • Float/double encoding β€” sign-bit handling differences between integer and IEEE 754
  • Multi-step mutations that aren't atomic β€” delete-then-insert, check-then-act
  • Concurrency contract violations β€” if the spec declares this construct thread-safe, check: are compound operations (get-or-create, check-then-act, position-then-read) actually atomic? If the spec declares not-thread-safe, note it but don't generate false concurrency findings
  • Switch/instanceof that don't cover all sealed interface subtypes
  • Silent truncation or Math.min instead of fail-fast on mismatched dimensions/sizes
  • Not-equals predicates interacting with null field values
  • Resource lifecycle β€” double-close, use-after-close, deferred exception aggregation
  • Validation that should happen at construction but is deferred to usage
  • Any patterns from adversarial KB entries loaded above

Level 2 β€” Inputs (who calls this construct, what do they pass?):

  • Are callers validated at the trust boundary, or is invalid input silently accepted?
  • Per project rules, should out-of-range values be rejected at entry rather than handled downstream? (fail-fast principle)
  • Can callers pass values that are technically valid but semantically wrong? (e.g., NaN as a score, negative capacity, inverted range bounds)

Level 3 β€” Outputs (what does this construct return, can consumers misuse it?):

  • Do returned references expose mutable internal state? (check accessors, not just constructors)
  • Are returned collections unmodifiable, or can consumers corrupt internal state?
  • Can the return value be in a state the consumer doesn't expect? (null, empty, partial)

Level 4 β€” Data carriers (records, DTOs, result types):

  • Do data carrier types enforce their own invariants at construction? (null fields, NaN scores, negative counts, empty required fields)
  • Do records with mutable fields (arrays, collections, MemorySegment) have correct equals/hashCode? (identity vs content semantics)
  • Are carriers immutable once constructed, or can state leak through accessors?

Scoping: Trace all 4 levels on every construct in the current work unit. If the work unit has many constructs, prioritize depth on constructs flagged by Lens A or Level 1 findings, but do not skip levels entirely.

Output

For each finding, note:

  • The construct and contract section it applies to
  • The finding type:
    • SPEC-REQ β€” directly operationalized from a spec requirement (Lens A, spec mode)
    • CONTRACT-GAP β€” gap not covered by any spec requirement (Lens A)
    • SPEC-BOUNDARY β€” spec-adjacent edge case (Lens B, spec mode)
    • SPEC-BLIND-SPOT β€” no spec coverage at all (Lens B, spec mode)
    • IMPL-RISK β€” standard implementation risk (Lens B)
    • DISPLACEMENT β€” behavior that must be verified as removed (from displacement resolution)
  • A specific defensive test case to add to the test plan
  • If from a spec requirement: the requirement ID (e.g., R3)
  • If from displacement: the displaced spec and requirement ID (e.g., displaced: F05.R3)

These findings feed directly into the test plan as a "Defensive (from spec analysis)" section. Do NOT write a separate spec-analysis.md file β€” the findings are integrated into the test plan in the next step.

Displacement verification

If the work-plan contains a ## Removal Work section, derive negative tests for each removal work unit. These verify that displaced behavior is gone:

For each removal entry (RW-1, RW-2, ...):

  1. Read the displaced spec requirement text to understand what behavior existed
  2. Derive a test that asserts the behavior no longer exists:
    • API removed β†’ calling it throws/returns an error
    • Format no longer supported β†’ input in old format is rejected with a clear error
    • Configuration removed β†’ using the old config produces a validation error
    • Behavioral change β†’ old behavior demonstrably not happening
  3. Tag as covers: <existing_id>.<req_id> (displaced) and finding type DISPLACEMENT

These tests go in the "Displacement verification" section of the test plan (between "Spec requirements" and "Defensive").


Step 2 β€” Write the test plan (in chat first)

Update status.md substage β†’ confirming-plan.

Coverage checklist (internal β€” apply before presenting the plan)

Before presenting the test plan, systematically check each construct against these categories. The refactor agent (step 2e) will check these exact categories later β€” gaps found there trigger an escalation cycle. Catching them here is much cheaper.

For each construct in the work plan:

CategoryWhat to testCommon gaps
Happy pathNormal operation with valid inputsRarely missed
Error conditionsEvery error case in the brief + contractMissing: errors not in brief but implied by types
Boundary valuesEmpty/zero, single element, max capacity, null/nilMost commonly missed β€” add at least one per construct
State transitionsInvalid state (e.g., use after close, double init)Missed when construct has lifecycle
ConcurrencyThread safety, ordering dependencies, race conditionsOnly if construct is shared/concurrent β€” skip if single-threaded
Type boundariesOverflow, underflow, precision loss, encoding limitsMissed for numeric types, byte buffers, serialization

The most commonly missed category is boundary values. For every construct, ask: "what happens at empty, at one, and at max?" If the answer isn't obvious from the contract, write a test for it.

Present the plan

Display:

── Test plan ───────────────────────────────────
TEST PLAN β€” <slug> (Cycle <n>)
─────────────────────────────────────────────────────────────
Happy path
  1. test_<name> β€” <scenario> β€” covers: <acceptance criterion>
  ...
Error and edge cases
  N. test_<name> β€” <scenario>
  ...
Boundary values
  N. test_<name> β€” <scenario>
  ...
Structural (from interface analysis)
  N. test_<name> β€” <scenario> β€” pattern: <round-trip | lifecycle | interaction | ...>
  ...
Spec requirements (from hardened specs)          ← only when specs loaded
  N. test_<name> β€” <scenario> β€” covers: <spec-id>.R<N>
  ...
Displacement verification (removal tests)       ← only when removal work exists
  N. test_<name> β€” <scenario> β€” covers: <spec-id>.R<N> (displaced)
  ...
Defensive (from spec analysis)
  N. test_<name> β€” <scenario> β€” finding: <CONTRACT-GAP | SPEC-BOUNDARY | SPEC-BLIND-SPOT | IMPL-RISK>: <description>
  ...
───────────────────────────────────────────────
Spec analysis: <N> SPEC-REQ + <N> CONTRACT-GAP + <N> IMPL-RISK findings β†’ <N> defensive tests
              [if specs loaded: <N> requirements operationalized, <N> untestable]
Does this cover the acceptance criteria? Any to add or remove?
───────────────────────────────────────────────

The plan MUST include "Boundary values", "Structural", and "Defensive" sections. If any section has no applicable tests (rare), state why explicitly. The structural section should list which interface patterns were checked and found not applicable. The defensive section should summarize how many findings came from each lens and any adversarial KB patterns that were checked.

When specs were loaded: The defensive section should additionally report:

  • How many spec requirements were operationalized (SPEC-REQ count)
  • How many requirements were marked UNTESTABLE and why
  • How many SPEC-BOUNDARY and SPEC-BLIND-SPOT findings were identified
  • Which requirement IDs each defensive test covers

Wait for confirmation. Update status.md substage β†’ writing-tests after confirm.


Step 3 β€” Write test files

Write to the test directory from project-config.md.

Idempotent: check whether each test already exists before writing. Append to existing test files rather than overwriting; do not duplicate tests. Before editing an existing test file, re-read it to pick up any additions from prior test-writing passes β€” stale reads cause Edit old_string mismatches or silent overwrites. After writing each test method, re-read the file to verify the method is present.

Rules:

  • Test names describe behaviour: test_returns_error_when_input_is_empty
  • Public interface only β€” no reaching into implementation details
  • Mocks/fakes for all external dependencies
  • Every test: arrange / act / assert clearly separated
  • Every construct MUST have at least one boundary value test (empty, zero, max, null/nil, single element) β€” the refactor agent checks for these and will escalate if missing
  • Every test MUST carry an @spec annotation pointing at the requirement it satisfies β€” see rules/spec-annotation-protocol.md for the format. Use the test plan's covers: <spec-id>.R<n> line as the write-time source: translate it directly into the test file's comment syntax above the test method. Skip only when no spec was loaded for this feature (the coverage table will be vacuous in that case).

Update status.md substage β†’ verifying-failures after writing.


Step 3a β€” Update spec coverage (mandatory if specs were loaded)

Run the coverage updater so the Test column of .feature/<slug>/spec-coverage.md reflects the annotations you just wrote. This is the structural enforcement of the annotation protocol β€” without it, /feature-pr cannot tell which requirements are bound.

if [[ -f .feature/<slug>/spec-coverage.md ]]; then
  bash .claude/scripts/spec-coverage.sh update \
    .feature/<slug>/spec-coverage.md --all
fi

Exit precondition. Read back the coverage file. Every row whose covers: entry in the test plan named a requirement you just wrote a test for MUST now show a non-pending value in the Test column. If any does not:

  1. Open the test file and confirm the @spec line is present, formatted correctly, and references the right <spec-id>.R<n>.
  2. Re-run the updater.
  3. If the cell still says pending, the annotation is malformed β€” spec-trace.sh did not match it. Fix the annotation and repeat.

Do not advance to Step 4 until every test row that should be annotated is annotated. The pipeline relies on this β€” partial annotation here becomes a silent gate failure at PR time.


Step 4 β€” Verify tests fail

Run the test suite (5-minute Bash timeout per tdd-protocol). If the suite times out: run individual test methods to isolate which test is hanging. For hanging tests, add a @Timeout annotation or rewrite with a non-blocking approach. Do not retry the full suite without isolating first.

Expected: all new tests fail with NotImplementedError or import/compile error.

If a test PASSES unexpectedly: investigate β€” stub may already be implemented, or test may not be testing the right thing. Do not hand off with passing tests at this stage.

Capture the full test runner output.


Escalation Review (--escalation only)

Entered when the Code Writer escalates a contract conflict via --escalation.

Step E1 β€” Load the escalation

Read the most recent code-escalation entry from cycle-log.md. Extract:

  • The test name and file
  • What the test expects
  • The constraint from the work plan
  • The conflict description
  • The escalation count (N of 3)

Read the test file and the relevant contract section from work-plan.md.

Step E1a β€” Check for spec conflict

If the escalation entry's conflict description contains "SPEC CONFLICT" or the substage in status.md is spec-conflict-detected, this is a requirement contradiction β€” not a test or contract problem.

Do NOT rewrite the test. Instead:

  1. Mark both the passing and failing tests as BLOCKED in cycle-log.md:

    ## <YYYY-MM-DD> β€” tests-blocked-spec-conflict
    **Agent:** πŸ§ͺ Test Writer
    **Cycle:** <n>
    **Blocked tests:** `<passing test>`, `<failing test>`
    **Reason:** Contradictory spec requirements β€” <R_N> vs <R_N>
    **Resolution:** Requires /spec-author to reconcile conflicting requirements
    ---
    
  2. Display:

    πŸ§ͺ TEST WRITER Β· spec conflict Β· <slug>
    ───────────────────────────────────────────────
    This escalation is a spec conflict, not a test or contract problem.
    Both tests are correct given their respective requirements β€” the
    requirements themselves contradict each other.
    
    Blocked tests:
      - <passing test> (covers: <R_N>)
      - <failing test> (covers: <R_N>)
    
    Run /spec-author to resolve the conflicting requirements, then
    re-run /feature-test "<slug>" to unblock.
    
  3. Stop. Do not proceed to Step E2.

Step E2 β€” Diagnose

Determine which of three cases applies:

  1. Test is wrong β€” the test asserts something not implied by the contract. Fix the test to match the contract. Run it to confirm it fails for the right reason (not-implemented, not wrong-assertion).

  2. Test is right, contract is ambiguous β€” the contract can be read multiple ways. Clarify the test (add a comment explaining intent) and adjust the assertion if needed. The Code Writer will re-read the test on next run.

  3. Contract itself is wrong β€” the work plan constraint contradicts the brief or an ADR, or is internally inconsistent. The Test Writer cannot fix this. Escalate to the Work Planner β€” see Step E3.

For cases 1 and 2, after fixing:

Append test-escalation-resolved to cycle-log.md:

## <YYYY-MM-DD> β€” test-escalation-resolved
**Agent:** πŸ§ͺ Test Writer
**Cycle:** <n>
**Test:** `<test name>` in `<file>`
**Diagnosis:** <test was wrong | contract was ambiguous>
**Fix:** <what changed>
**Escalation count:** <N> of 3
---

Update status.md substage β†’ escalation-resolved.

Display:

πŸ§ͺ TEST WRITER Β· escalation resolved Β· <slug>
───────────────────────────────────────────────
Test: <test name>
Diagnosis: <test was wrong | contract was ambiguous>
Fix: <what changed>

Re-run implementation:
  /feature-implement "<slug>"<  --unit WU-<n>>

Stop.

Step E3 β€” Escalate to Work Planner

Before escalating, check the escalation counter.

Read cycle-log.md and count test-to-planner-escalation entries for the same contract/construct.

3rd escalation on the same contract: hard stop. Do NOT escalate to the Work Planner again. Instead:

πŸ›‘  ESCALATION LIMIT Β· Test Writer β†’ Manual Resolution
───────────────────────────────────────────────
The same contract issue has been escalated 3 times without resolution.
Contract: <construct name>
Work plan section: <reference>

Automatic resolution is not working. Please review the conflict manually:
  1. Check the contract in work-plan.md
  2. Check the acceptance criteria in brief.md
  3. Check any governing ADRs
  4. Fix the work plan, then re-run /feature-test "<slug>"

If the brief itself is wrong, revisit /feature "<slug>".

Update status.md substage β†’ escalation-limit-reached. Stop.

Under the limit: proceed with escalation.

Append test-to-planner-escalation to cycle-log.md:

## <YYYY-MM-DD> β€” test-to-planner-escalation
**Agent:** πŸ§ͺ Test Writer
**Cycle:** <n>
**Contract:** <construct name>
**Work plan section:** <reference>
**Conflict:** <what the contract says vs. what it should say>
**Brief reference:** <acceptance criterion that contradicts the contract>
**Escalation count:** <N> of 3
---

Update status.md substage β†’ escalated-to-work-planner.

Display:

⚠️  ESCALATION Β· Test Writer β†’ Work Planner  (<N>/3)
───────────────────────────────────────────────
Contract conflict β€” the work plan constraint cannot satisfy the brief.
Contract: <construct name>
Problem: <paragraph>
Brief reference: <acceptance criterion>

The Work Planner needs to revise this contract. Run:
  /feature-plan "<slug>"
Then re-run the test β†’ implement cycle for this construct.

Stop.


Step 5 β€” Write test-plan.md and log

Write .feature/<slug>/test-plan.md (or append cycle section if it exists):

## Cycle <n> β€” <YYYY-MM-DD>

### Tests Written

| Test name | File | Construct | Acceptance criterion |
|-----------|------|-----------|---------------------|
| test_<n> | <path> | <construct> | <criterion> |

### Failure output (expected)
<test runner output> ```

Coverage

  • Acceptance criteria covered: <n>/<total>
  • Error cases covered: <n>
  • Spec requirements operationalized: <n>/<total> (omit if no specs loaded)
  • Untestable requirements: <list or none> (omit if no specs loaded)
  • Gaps noted: <any>

Update status.md:
- Testing cycle N β†’ `complete`
- TDD Cycle Tracker: Tests written β†’ today
- substage β†’ "tests verified failing"
- Stage Completion table: Testing row β†’ Est. Tokens `~<N>K` (project-config ~1K +
  brief ~2K + work-plan section ~2K + test files written)

Append `tests-written` entry to cycle-log.md.
Update `.feature/CLAUDE.md`.

---

## Step 6 β€” Hand off

Read `automation_mode` from status.md.

### Determine next stage: hardening or implement

Check the work-plan contracts for domain lens signals to decide whether
hardening is needed. Count constructs with any of these contract properties:
- Closeable/AutoCloseable, close/cleanup mentions β†’ resource_lifecycle
- Cross-module dependencies β†’ contract_boundaries
- Thread-safety mentions, `shares_state` edges β†’ concurrency
- Encode/decode, serialize/deserialize β†’ data_transformation
- Mutable state shared by 2+ constructs β†’ shared_state

**If zero lens signals OR only 1 construct in the work plan:**
- Skip hardening β€” chain directly to `/feature-implement`.

**If 1-2 lens signals AND 2-5 constructs:**
- Chain to `/feature-harden "<slug>" --lite<  --unit WU-<n>>`.

**If 3+ lens signals OR 6+ constructs:**
- Chain to `/feature-harden "<slug>"<  --unit WU-<n>>`.

### Chain execution

**If `automation_mode: autonomous`:**

Display the summary then chain immediately without prompting:

─────────────────────────────────────────────── πŸ§ͺ TEST WRITER complete Β· <slug> Β· Cycle <n>< Β· WU-<n>> Tokens : <TOKEN_USAGE> ─────────────────────────────────────────────── Tests written and verified failing. Cycle <n>< Β· WU-<n>>.

Starting <hardening | implementation> Β· type stop to pause ───────────────────────────────────────────────


Then invoke the next stage as a sub-agent immediately.

**If `automation_mode: manual` (or not set):**

Display:

─────────────────────────────────────────────── πŸ§ͺ TEST WRITER complete Β· <slug> Β· Cycle <n>< Β· WU-<n>> Tokens : <TOKEN_USAGE> ─────────────────────────────────────────────── Tests written and verified failing. Cycle <n>< Β· WU-<n>>.


Use AskUserQuestion with two options:
- "Continue"
- "Stop"

If "Continue": invoke the next stage as a sub-agent immediately.
If "Stop":

When you're ready: /feature-harden "<slug>"< --unit WU-<n>> (or /feature-implement if skipping)

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,758. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.