Plan milestone
Skill ronniepinnell/casper/collection/planning-and-lifecycle/plan-milestone
π» The friendly ghost in your git. Your AI said done β Casper makes it prove it. Claim-evidence hooks + a verdict ledger for Claude Code.
npx -y skills add ronniepinnell/casper --skill plan-milestoneAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 27 days oldThe repository was created 27 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Break a milestone goal into epics and tasks with dependency ordering, acceptance tests, and task manager ticket creation. Works with any task manager via the adapter layer.
SKILL.md
97.6 KB, ~25.4k tokens by cl100k_base, as published. Nobody here has run it
Factory mode: If
CLAUDE_AUTOis set in the environment, skip these instructions and instead run./{scripts_dir}/skills/plan-milestone.sh "$@"via Bash β the shell script handles autonomous dispatch.
MCP Tool Map (Gemini/Codex): See
.claude/skills/_shared/mcp-tool-map.mdfor tool name equivalents.
Plan Milestone β Epic & Task Breakdown
Takes a milestone goal and produces a full execution plan: epics with dependency ordering, tasks with test criteria, and task manager tickets ready for the factory.
Milestone Naming Rule β HARD GATE
Milestone titles must describe a CAPABILITY being added, not a STATE being reached.
The recurring failure mode is the "DAT004β008" pattern: milestones named
Finish the Schema, Complete Phase 10, Wrap up the cleanup are open-ended
buckets that absorb scope until they're marked Done without ever shipping a
discrete capability. Capability names force a clear acceptance test
(Add Lineage Tracking β "lineage column exists, populated for all rows").
Rejected patterns (regex enforced before milestone creation)
| Reject if title matches | Why |
|---|---|
^(Finish|Complete|Close|Wrap[- ]up|Finalize)\b | "Finish X" is state, not capability |
^Final\b | "Final pass" never ships a thing |
substring: the schema, the cleanup, the audit, remaining work | bucket nouns absorb scope |
Canonical regex (case-insensitive β apply with (?i) inline flag or
/.../i in JS-style engines). Markdown escapes the pipes in the table
above; the raw patterns are:
^(Finish|Complete|Close|Wrap[- ]up|Finalize)\b
^Final\b
Plus literal substring checks (case-insensitive) for:
the schema, the cleanup, the audit, remaining work.
Required pattern
Title MUST start with an action verb describing the capability added:
Add Β· Build Β· Enable Β· Ship Β· Deploy Β· Migrate Β· Lock-in Β·
Harden Β· Retire Β· Replace Β· Extract Β· Introduce Β· Wire
| Bad (state) | Good (capability) |
|---|---|
Finish the Schema | Add Lineage Tracking |
Complete Phase 10 | Ship Multi-Tenant RLS |
Wrap up the cleanup | Retire Legacy CSV Loader |
Final auth pass | Harden JWT Refresh Flow |
The skill MUST refuse to create a milestone whose title matches a rejected
pattern. Surface the regex hit, suggest the verb list, and force the operator
to rename before save_milestone is called.
Step 0: Load Project Context (MANDATORY)
Before any planning, read .claude/project-context.md to get:
task_managerβ which adapter to usetask_prefixβ ticket prefix (e.g.APP,BEN,WEB)task_team_idβ team/project identifiermilestone_nounβ what this project calls milestonesmain_branchβ PR target branchareasβ project area codes (if defined)storage_backendβ load.claude/skills/_shared/storage/{storage_backend}.md; route all memory reads/writes (record_decision,list_decisions,save_task_plan,record_test_run, β¦) through its operations. Absent βnone(reads empty, writes discarded).storage_schema_intel/storage_schema_qaname the namespaces (defaultsintel/qa).agentsβ capabilityβagent map; reference agents below as{agents.<capability>}. Absent β template defaults (completion_audit: completion-audit,spec_audit: spec-audit,pragmatism_audit: pragmatism-audit,scalability_audit: future-self). If a mapped agent is unavailable, skip that step with a logged note.
Then load the task manager adapter:
Read.claude/skills/_shared/adapters/{task_manager}.md
If .claude/project-context.md is missing: stop and tell the user to run /project-init first.
GATE ENFORCEMENT: This skill is the ONLY path to write task manager issues. A PreToolUse hook (
guard-task-writes.sh) blocks task creation calls unless.claude/.plan-milestone-activeexists. This skill creates that flag on entry and removes it on exit.Step 0b (MANDATORY β before any planning):
touch.claude/.plan-milestone-activeFinal step (MANDATORY β after all issues created):
rm -f.claude/.plan-milestone-active
Naming Conventions (CRITICAL)
Milestone IDs β {AREA_CODE}{3D}{LETTER}
Format: 3-letter area code + 3-digit sequence + letter suffix.
Area codes are defined per project in .claude/project-context.md under areas:.
If no areas are defined, auto-derive the code from the area name:
Auto-derivation rules (applied in order):
1. Split area name into words
2. If multi-word: take first letter of each word, uppercase, max 3 chars
"Computer Vision" β CV β pad to CVX? No β use first 3 letters of each word initial: CVX
"Machine Learning" β ML β MLX (pad single/double initials to 3 with X)
"Analytics" β ANL (first 3 consonant-rich chars)
3. If single word β€ 3 chars: uppercase as-is
4. If single word > 3 chars: first 3 chars, uppercase
"Analytics" β ANL | "Factory" β FCT | "Platform" β PLT
5. Confirm with user before creating first milestone in a new area
Sequence is auto-incremented: fetch existing milestones for this area from the task manager, find the highest sequence number, add 1. Start at 001 for new areas.
Examples (after derivation):
Computer Vision, seq 5, phase A β CVX005A
Machine Learning, seq 4, phase A β MLX004A
Factory, seq 3, phase A β FCT003A
Z suffix = hardening/verification milestone (e.g. CVX005Z β Verify CV Pipeline E2E).
NEVER use decimal suffixes. Increment the letter (AβBβCβ¦) for sub-phases.
Epic Titles β [{MILESTONE_ID}] {Title}
[MLX004A] V3 Hybrid Architecture
[FCT003A] Budget & Safety Gates
[CVX005A] VERIFY: Baseline Verification
Task Titles β plain descriptive
No prefix needed β the parent relationship provides context.
Milestone Labels
Every epic and task gets a milestone:{CODE###X} label for cross-project filtering.
Git Branching β Derived from type:* Label
Each epic gets ONE branch. All tasks commit to that branch. ONE PR per epic to {main_branch}.
type:* label | Branch prefix | Example |
|---|---|---|
type:feature | feature/ | feature/{prefix}-1371-v3-architecture |
type:fix | fix/ | fix/{prefix}-993-auth-regression |
type:refactor | refactor/ | refactor/{prefix}-997-split-module |
type:test | test/ | test/{prefix}-1000-period-validation |
type:audit | audit/ | audit/{prefix}-996-verify-m0004d |
type:docs | docs/ | docs/{prefix}-1010-doc-automation |
{prefix} = task_prefix from project-context.md (e.g. BEN, APP, WEB).
Branch format: {type}/{prefix}-{epic_id}-{short-kebab-name}
PR target: {main_branch} from project-context.md (NEVER main if configured otherwise)
Commit prefix: [{TYPE}] {prefix}-{task_id}: {what changed}
Invocation
/plan-milestone "Computer Vision" 5 "Ship tracking pipeline"
/plan-milestone CVX005E --update # patch existing issues to match current spec
/plan-milestone CVX005E --update --dry-run
If a milestone ID is passed directly (e.g. CVX005E), skip code derivation and use as-is.
When called from /ccb, the --from-ccb flag is implicit and CCB context (blockers, debt, acceptance criteria) is already in conversation.
--update Flag β Bring Existing Issues Up to Date
When --update is passed, skip Steps 2-4 (no new planning). Instead:
Update Step 1: Load All Existing Issues
issues = list_issues(team: "{task_team_id}", query: "[{milestone_id}]")
Update Step 2: Resolve Correct Project & Milestone IDs
Same as Step 5a-pre β look up the Linear project by code prefix and the milestone by name.
Update Step 3: Audit Each Issue
For every epic and task in this milestone, check:
| Field | Check | Fix |
|---|---|---|
projectId | Matches resolved project? | Set it |
milestoneId | Matches resolved milestone? | Set it |
labels | Has milestone:{CODE} label? | Add it |
labels | Has model:* label? (epics + tasks) | Add default from routing table |
labels | Has machine:* label? | Add default from routing table |
labels | Has type:* label? (epics) | Add based on title prefix |
labels | Has epic or task label? | Add based on parentId |
| Description | Has ## Execution Context? | Add from routing defaults |
| Description | Has ## Goal with real content? | Flag as thin β needs manual fill |
blockedBy | Set correctly per dependency graph? | Flag if missing |
| Parent | Tasks have correct parentId? | Flag if orphaned |
Update Step 4: Show Audit Report
PLAN-MILESTONE UPDATE: {milestone_id}
ββββββββββββββββββββββββββββββββββββ
Issue Field Current Fix
ββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
{prefix}-3433 projectId (none) β Computer Vision
{prefix}-3433 milestoneId (none) β CVX005E
{prefix}-3434 labels missing model:* β model:sonnet
{prefix}-3435 Execution Context (missing section) β Add default block
{prefix}-3436 milestone label (none) β milestone:CVX005E
{fix_count} fixes needed across {issue_count} issues
If --dry-run: print report and stop.
If not --dry-run: use AskUserQuestion to confirm, then apply all fixes via save_issue.
Update Step 5: Verify
Run the same verification sweep as Step 5g β confirm all issues are now linked correctly.
Step 0b: Operator Alignment Check
Ask before loading specs or generating anything:
Before planning {milestone_id}: anything to discuss, clarify, or flag?
(scope changes, locked decisions, timeline constraints, things that changed)
Use AskUserQuestion with options:
- Nothing β let's plan (default)
- I have something to flag β free-text; capture decisions via
record_decision(storage) then continue - Hold β scope isn't settled yet β STOP, do not proceed
If CLAUDE_AUTO=1: skip this step.
Grill me (strategic milestones): When the milestone has β₯ 3 epics planned or touches schema/ETL/auth, ask 2β3 clarifying questions before generating the epic breakdown:
- "What does 'done' look like for this milestone β what can the Operator do that they can't do today?"
- "Are there any scope items that should NOT be included, even if they seem related?"
- "Any external dependencies (integrations, data sources, vendors) that could block this?"
Capture answers as decisions if they constrain scope. Skip if Operator answers "looks good" to any question.
Step 0c: Honesty Stack Mode Selection
Before generating epics, lock in the honesty-stack mode that the milestone and every child epic will inherit. The mode controls which honesty hooks fire at /epic close / /milestone close:
fullβ all hooks (R-43 lint + audit-doubt + verifier-isolation + transition + dependency-audit). Use for P0 milestones (FCT*, DAT*, HARDEN, Z) where the cost of a missed honesty signal is high.liteβ cheap hooks only (R-43 lint + transition-validator). The default β keeps routine epics cheap.offβ nothing fires. Emergency override for trivial fixes / Operator directive.
Resolution order (highest precedence first)
HONESTY_MODEenv var (non-interactive fallback for autonomous runs).- The Operator's
AskUserQuestionchoice in this step. resolve_honesty_mode({milestone_id})fromscripts.factory.lifecycle_helpersβ applies theper_milestone_overridesglobs declared in.claude/project-context.md.
Procedure
import os
from scripts.skills.plan_milestone_helpers import (
read_honesty_mode_from_description,
record_honesty_mode_decision,
resolve_honesty_mode,
upsert_honesty_block,
)
resolved_default = resolve_honesty_mode(milestone_id) # e.g. "full" for FCT*
env_override = (os.getenv("HONESTY_MODE") or "").strip().lower()
existing_mode = read_honesty_mode_from_description(current_milestone_description)
if env_override in {"full", "lite", "off"}:
# CLI override always wins, even over a persisted block.
chosen_mode = env_override
rationale = "HONESTY_MODE env var override"
elif existing_mode in {"full", "lite", "off"} and not honesty_block_is_stale(current_milestone_description):
# CR PR #9817: re-running /plan-milestone on an already-planned milestone
# must NOT re-prompt the operator for the mode. Reuse the persisted
# block when it's fresh.
chosen_mode = existing_mode
rationale = "Reused persisted honesty mode from milestone description"
elif os.getenv("CLAUDE_AUTO") == "1":
chosen_mode = resolved_default
rationale = f"CLAUDE_AUTO=1, using resolved default for {milestone_id}"
else:
# Interactive: surface the resolved default + ask the operator.
# Use AskUserQuestion with three options. The label of the resolved
# default carries "(Recommended)".
chosen_mode = ask_via_AskUserQuestion(resolved_default)
rationale = f"Operator choice (resolved default was {resolved_default})"
The AskUserQuestion block:
Question: "Honesty stack mode for {milestone_id}? (resolved default: {resolved_default})"
Options:
- "{resolved_default} (Recommended)" β recap of what {resolved_default} runs
- "{other_mode_1}" β recap of what {other_mode_1} runs
- "{other_mode_2}" β recap of what {other_mode_2} runs
Persistence
After resolving chosen_mode:
# 1. Write the block into the milestone description (idempotent β replaces
# any existing block, otherwise inserts at end).
new_description = upsert_honesty_block(
description=current_milestone_description,
mode=chosen_mode,
milestone_id=milestone_id,
rationale=rationale,
)
# Apply via save_milestone (or save_issue for the
# tracker proxy).
# 2. Record the choice via record_decision (supabase β {storage_schema_intel}.decisions). Returns the decision_number on
# success, None on failure (never raises).
record_honesty_mode_decision(milestone_id, chosen_mode, rationale)
The block rendered by upsert_honesty_block looks like:
<!-- honesty-stack:begin -->
## Honesty Stack
- mode: `full`
- resolved: 2026-05-18 by /plan-milestone for FCT011C
- rationale: Operator choice (resolved default was full)
<!-- honesty-stack:end -->
/milestone start, /pipeline run, and /epic start read this block via read_honesty_mode_from_description() to seed each session's honesty_mode column.
Skip conditions
CLAUDE_AUTO=1β use resolved default silently (no prompt).- Milestone description already contains a non-stale honesty-stack block AND no CLI override β re-use the existing mode, do NOT re-prompt.
Step 1: Spec Grounding β HARD GATE (do NOT skip, do NOT plan from memory)
Planning from the milestone title or from memory β instead of from the spec files on disk β is the #1 cause of drift and false Done marking. You may not write acceptance criteria (Step 2), epics, or tasks until you have produced and shown the Spec Grounding Digest below. This gate is as binding as the naming gate above.
1a. Discover the relevant specs β run these, don't guess
ls {spec_dir}/ ; cat {spec_dir}/.abstract.md 2>/dev/null # what specs exist
# milestone-goal keywords β matching specs (substitute 2-4 real keywords):
grep -rilE '<keyword1>|<keyword2>|<keyword3>' {spec_dir}/ rules/ | head -20
test -f {spec_dir}/GITHUB_FACTORY_SPEC.md && echo "factory spec present"
Also load the constraints that bound the plan:
list_decisions()β storage op (supabase β{storage_schema_intel}.decisions) β locked decisions (D-block etc.) that constrain choices.- Prior-milestone
**Spec Section:**refs and any open carry-over. - Open / recently-closed issues:
gh issue list --state open --json number,title,labels,milestone --limit 100.
1a-verify. External model cross-check β did you miss any specs? (MANDATORY)
Before reading anything, send your discovered spec list + the milestone goal to Ollama or Gemini. The external model sees only filenames and the goal β it cannot read the files. Its job is to flag specs you might have missed based on filename pattern and milestone topic alone.
python3 {scripts_dir}/skills/verify_spec_coverage.py \
--goal "{milestone goal as plain English}" \
--found-specs "{spec_dir}/FOO.md {spec_dir}/BAR.md rules/areas/etl.md" \
--spec-dirs "docs/specs docs/runbooks rules/areas"
The script calls Ollama (β Gemini fallback) and returns:
- Confirmed: specs your search found that the model agrees are relevant
- Missed: specs in the tree the model thinks you should also read
- Irrelevant: specs your search found that the model thinks don't apply
Any file in Missed is a mandatory read β add it to your list before 1b. If the external model is unavailable, skip with a logged warning and proceed.
1b. READ each candidate spec end-to-end, then write the digest
Open every spec relevant to this milestone with the Read tool and read it in full. Then
write a Spec Grounding Digest to .claude/.plan-specs-{milestone_id}.md and show it
to the Operator. One row per relevant spec section:
Spec ref (file#Β§) | Lines read | Requirement (VERBATIM quote from the file) | Status today | Epic |
|---|---|---|---|---|
{spec_dir}/X.md#Β§4.2 | L120β138 | "the manifest MUST declare every generated tableβ¦" | none | E01 |
{spec_dir}/X.md#Β§4.3 | L139β151 | "CI fails if declared β reality" | partial | E02 |
The verbatim quote column is the forcing function: you cannot fill it without having
opened the file. Paraphrase-only rows are rejected. Every quote needs its LstartβLend.
1c. Spec Quiz β mechanical verification (do NOT skip)
After writing the digest, run the quiz against every spec file you read. The quiz pulls verbatim lines and asks you to fill them in β it cannot be passed without having read the file.
python3 {scripts_dir}/skills/spec_quiz.py <spec_file1> [<spec_file2>...] --questions 2
- Pass (β₯70%): proceed to gate check.
- Fail: return to 1b, re-read the flagged file end-to-end, retry once.
- Still fail: stop. Tell the Operator which spec is unclear before continuing.
Also verify the digest structure is valid (catches empty-quote cells):
python3 {scripts_dir}/skills/spec_quiz.py --digest.claude/.plan-specs-{milestone_id}.md
This must return OK before Step 2.
1d. Gate check β ALL must be true before Step 2
- Spec quiz passed (β₯70%) for every spec file read.
- Digest structure check returned
OK(no empty quote cells). - Every spec path in the digest passed
test -f(it is a real file). - Every row has a verbatim quote and a line range β no empty/paraphrased quote cells.
- Every milestone epic in Step 3 maps to β₯1 digest row via its
**Spec Section:**line. - The digest file
.claude/.plan-specs-{milestone_id}.mdexists and is non-empty. - Spec gaps are explicit: each section is marked
none/partial/doneβ gaps drive the epic breakdown.
If you cannot quote a requirement, you have NOT read that spec β return to 1b. These rows are the source feeding the Β§46 spec-coverage producer hook (Step 5a-post); a missing digest there means you skipped this gate.
Step 2: Generate Acceptance Criteria
Before any epic planning, define milestone-level acceptance criteria. These are the tests that the hardening epic will run.
Format:
## {milestone_id} Acceptance Criteria
### Functional
- [ ] {User-visible behavior that must work end-to-end}
- [ ] {Another behavior}
### Technical
- [ ] {Performance/reliability requirement}
- [ ] {Integration requirement}
### Quality
- [ ] All new code has test coverage
- [ ] No new files over 500 lines
- [ ] All spec sections for this milestone have implementation
OUTCOME-BASED CRITERIA RULE (CRITICAL): Every acceptance criterion MUST assert an observable OUTCOME, never an operation.
| BANNED (operation-based) | REQUIRED (outcome-based) |
|---|---|
| "Migration X applied successfully" | "Table X has columns A (bigint), B (text), C (uuid)" |
| "Script ran without errors" | "Query SELECT count(*) FROM X returns > 0" |
| "File was created" | "File exports function Y and passes type-check" |
| "DROP COLUMN ran" | "Column Z does NOT exist in information_schema" |
| "RLS policy created" | "pg_class.relrowsecurity = true for table X" |
Operation-based criteria are how false Done marking happens β a DROP COLUMN IF EXISTS silently no-ops, the agent checks "did it run?" (yes), marks Done, and the column is still there. Outcome-based criteria catch this because they check the actual state.
Present to Operator for approval. Adjust based on feedback.
Step 2b: Write Epic Test Files BEFORE Tasks (TDD β Red First, Independent Model)
HARD RULE: Claude does NOT write the tests. A separate model writes them. Claude's implementation plan is NOT shown to the test writer β only the Spec Grounding Digest and Exit Criteria. This is the only way tests are independent. Self-review is not review.
Step 2b-0: Dispatch Test Writing to External Model (MANDATORY)
After Step 2a (Exit Criteria are written), dispatch test writing to an external model. The test writer sees ONLY the spec digest and exit criteria β never Claude's internal planning notes.
Preferred dispatch order (try in sequence, use first available):
-
Ollama (local, free) β try
codestral:22bordeepseek-coder-v2first:ollama list 2>/dev/null | grep -E "codestral|deepseek-coder|qwen2.5-coder" | head -3If a suitable model is available:
# Write the test brief (spec + exit criteria ONLY β no implementation context) cat > /tmp/test_brief_{milestone_id}.md << 'EOF' # Test Writing Brief β {milestone_id} {epic_id} You are a grumpy, adversarial test writer. You do not trust the implementation. Your job: write failing (RED) pytest tests that prove the exit criteria below are met. You have NOT seen the implementation. Write tests that WOULD CATCH a lazy implementation. ## Spec (verbatim from spec digest) {paste relevant rows from.claude/.plan-specs-{milestone_id}.md} ## Exit Criteria to Test {paste this epic's Exit Criteria from Step 2a} ## Rules - Every criterion gets β₯1 test function - Fail with raise NotImplementedError("RED: {criterion}") β NEVER pytest.skip - Docstring states which criterion is proven - Assert OUTCOMES, not "did the function run" - Be adversarial: write the test that would catch a stub or fake EOF ollama run codestral:22b < /tmp/test_brief_{milestone_id}.md > /tmp/test_output_{milestone_id}.pyReview the output. If sensible, copy to the correct test path (see naming below). If garbage, fall back.
-
Codex (if available) β same brief, same isolation constraint. Codex does NOT see Claude's plan.
-
Claude as fallback (last resort only) β only if Ollama and Codex are both unavailable. If Claude must write its own tests, explicitly log in the milestone description:
β οΈ SELF-TESTED: No external model available at planning time. Tests written by Claude. Flag for adversarial review before epic closes.Then write tests using the Exit Criteria only β deliberately ignore the implementation plan.
For every epic in this milestone, write the actual test file(s) to disk NOW β before any task is created in Linear. These tests FAIL immediately (RED). They turn GREEN when the epic is done. This is not optional. If there is no test file, there is no proof of completion.
Tests are always specific to what THIS epic builds. Never write generic scaffolding or
pytest.skip placeholders. Derive the test content directly from the epic's Exit Criteria β
each criterion becomes one or more test functions.
Layer Routing Table β which suite for which epic type
| Epic builds... | Test layer | Suite folder | When it runs |
|---|---|---|---|
| A function, class, hook, CLI command | 1 β Unit | tests/unit/ | Every commit on epic branch |
| DB queries, API endpoints, cross-module flows | 1+2 β Integration + Contract | tests/integration/, tests/contract/ | Every PR |
| Schema migration, RLS policy, column change | 2+8 β Contract + Migration | tests/contract/, tests/migrations/ | VERIFY-1 gate |
| ETL pipeline output, data transformation | 6 β Canonical | tests/canonical/ | VERIFY-2 gate |
| Auth, RLS, COPPA/parental_consent, anon access | 5 β Security | tests/security/ | VERIFY-4 gate |
| Dashboard page, user flow, UI interaction | 3 β E2E | tests/e2e/ | VERIFY-4 gate |
| Fixes a known past bug or false Done claim | 8 β Regression | tests/regression/ | Every commit forever |
| Invariants that must hold for any input | 4 β Property | tests/property/ | VERIFY-FINAL + nightly |
| Latency SLO, throughput, scale | 7 β Performance | tests/performance/ | VERIFY-FINAL + nightly |
Most epics need Layer 1 (unit) + one domain-specific layer. Never skip Layer 1.
Naming Convention (mandatory)
tests/{suite}/test_{milestone_id_lower}_{epic_short}_{what}.py
Examples:
tests/unit/test_proc001a_e01_karen_blocks_done.py
tests/contract/test_dat008a_e03_event_schema.py
tests/security/test_proc001a_e02_apply_migration_blocked.py
tests/regression/test_ben3706_onb001a_invite_flow.py β issue-named regressions
How to write the test (not a placeholder)
Read the epic's Exit Criteria. For each criterion, write one test function that:
- Asserts the OUTCOME directly β never "did it run?" but "does the result match?"
- Fails right now with
raise NotImplementedError(f"RED: {what this epic must deliver}")β NOTpytest.skip.skiphides the test.NotImplementedErrorshows the gap. - Has a docstring stating which Exit Criterion it proves
# tests/unit/test_proc001a_e01_karen_blocks_done.py
"""
PROC001A E01: completion-audit pre-commit hook tests.
Written RED before implementation. All must pass before epic closes.
"""
import pytest
class TestKarenDoneGuard:
def test_blocks_commit_when_test_file_missing(self, tmp_branch):
"""Exit Criterion: git commit with 'Done' fails if ## Tests file doesn't exist."""
raise NotImplementedError("RED: guard-done-marking.sh not yet implemented")
def test_blocks_commit_when_test_file_has_zero_functions(self, tmp_branch):
"""Exit Criterion: commit blocked if listed test file has no test_ functions."""
raise NotImplementedError("RED: guard-done-marking.sh not yet implemented")
def test_allows_commit_when_tests_pass(self, tmp_branch, passing_test_file):
"""Exit Criterion: commit succeeds when listed tests all pass."""
raise NotImplementedError("RED: guard-done-marking.sh not yet implemented")
Three layers required per exit criterion (thin tests are rejected)
Every exit criterion must produce tests in all three layers. One criterion = three test functions minimum:
| Layer | Class prefix | What it asserts | Docstring must contain |
|---|---|---|---|
| 1 Mechanical | TestMech_ | Code runs, returns correct type/shape | "Mechanical:" |
| 2 Spec/Contract | TestSpec_ | Output matches verbatim spec requirement | "Spec: [quoted requirement]" |
| 3 Outcome | TestOutcome_ | End user sees correct result | "Outcome: As a [user]..." |
A test file with only Layer 1 tests will be rejected by the TDD gate. Layer 2 docstrings MUST quote the spec verbatim β no paraphrase.
After writing all epic test files
pytest tests/ --collect-only -q 2>/dev/null | tail -5 # confirm tests are collected
pytest tests/unit/test_{milestone}_{epic}*.py # confirm they FAIL (RED)
git add tests/
git commit -m "[TEST] {milestone_id}: Write RED test files for all epics (pre-implementation)"
This commit is the proof that TDD is real. The factory cannot claim an epic Done unless its test file exists and pytest returns 0.
The ## Tests section of every Linear task must reference these exact file paths.
Step 2b.1: Consume CCB Phase 3.26 output β testing-scenarios + walkthrough twin
If this milestone was planned via /ccb, Phase 3.26 already produced, per
feature epic: a ## Testing Scenarios block, a walkthrough-twin issue under
master_verification_milestone (.claude/project-context.md), and a row in
config/test_coverage_matrix.json. /plan-milestone does NOT regenerate
these β it writes them verbatim into the epic description alongside ## Tests, and confirms the twin link:
for each feature epic:
assert epic.body contains "## Testing Scenarios" # from CCB 3.26
assert epic.body contains "Verification twin: {id}" # walkthrough issue link
if either missing:
# This milestone was NOT run through /ccb (or /ccb predates Phase 3.26).
# /plan-milestone must generate them itself β do not skip.
build the testing-scenarios block per the template in
`.claude/skills/ccb/SKILL.md` Β§ Phase 3.26, using this epic's acceptance
tests + spec refs as source material.
create the walkthrough-twin issue under master_verification_milestone.
append the row to config/test_coverage_matrix.json.
This is what makes planning output "feature epic AND its verification
twin" with zero extra Operator steps β the twin either arrives pre-built from
/ccb or gets built here, but it always exists before the epic is written
to the task manager.
Step 2c: Write Acceptance Test Scaffolding (Outcome Verification β Red First)
For every epic, also create acceptance test scaffolding in tests/acceptance/ for each applicable tier. These answer "does this actually work for a real user?" β not just "does the code exist?"
Tier selection β auto-determine from what the epic builds:
| Epic builds | Scaffold file | Key assertion to pre-write |
|---|---|---|
| CLI command | tests/acceptance/cli/test_{milestone}_{epic}.py | "Running {project_cli} X Y returns expected output against live DB" |
| API endpoint | tests/acceptance/api/test_{milestone}_{epic}.py | "GET/POST to /api/X returns schema matching TS interface" |
| Dashboard page/component | tests/acceptance/ui/test_{milestone}_{epic}.py | "Page loads, key data renders, primary action works" |
| ETL table / migration | tests/acceptance/data/test_{milestone}_{epic}.py | "Table exists, row count in range, idempotency holds" |
All acceptance tests must:
- Use
pytest.mark.skipif(not os.environ.get("SUPABASE_URL"), reason="requires live credentials") - Fail RED with
raise NotImplementedError("RED: {what this must prove}")until the epic is built - Have a docstring stating the user story: "As a {user}, I can {action} so that {outcome}"
Also write the User Story and End-User Verification stub for each epic directly into the Linear epic description's ## End-User Verification section. Pre-fill the format:
## End-User Verification
**As a {user type}, you can now:** {action in plain English}
**To verify:**
1. {Step 1 β e.g. "Run `{project_cli} X` or go to /page/url"}
2. {Step 2 β what to look for}
3. Expected: {what correct looks like}
**Regression check:** verify {top 2 adjacent features} still work.
Write TBD β fill at close for the regression check if unknown at planning time.
After writing all acceptance scaffolding:
pytest tests/acceptance/ --collect-only -q 2>/dev/null | tail -5 # confirm collected
git add tests/acceptance/
git commit -m "[TEST] {milestone_id}: Write acceptance test scaffolding (RED)"
Step 3: Epic Breakdown
3a: Identify Epics
Group related work into epics. Each epic should:
- Map to 1-2 spec sections
- Be completable in 1-3 days of factory time (5-15 issues)
- Have clear entry criteria (what must exist before this epic starts)
- Have clear exit criteria (how do we know this epic is done)
3b: Dependency Ordering
Build a dependency graph:
[FCT003A] Codex/Cursor CLI β no deps
[FCT003A] Clone Slots β depends on CLI epic
[FCT003A] Review Loop β depends on Clone Slots
[FCT003A] Auto-Merge β depends on Review Loop
[FCT003Z] HARDEN β depends on ALL above
3b.1: blockedBy Convention (CRITICAL β for all agents)
Every blockedBy relationship must be set explicitly via Linear. Never rely on implied ordering.
Within-Milestone Patterns
VERIFY epic: blockedBy = [] (always first, no deps)
Feature epic: blockedBy = [VERIFY] (all feature epics blocked by VERIFY)
Feature epic (cross-dep): blockedBy = [VERIFY, {prefix}-{other_feature}]
VERIFY-MECH epic: blockedBy = [ALL feature epics] (if selected in Step 3d.5)
HARDEN epic: blockedBy = [VERIFY-MECH] (or [ALL feature epics] if no VERIFY-MECH)
VERIFY-HUMAN epic: blockedBy = [HARDEN] (if selected in Step 3d.5)
Standard within-milestone blockedBy call:
# VERIFY: no blockers
# Feature epic E02:
save_issue({ id: "{E02}", blockedBy: ["{prefix}-{VERIFY}"] })
# Feature epic E03 (also blocked by E02):
save_issue({ id: "{E03}", blockedBy: ["{prefix}-{VERIFY}", "{prefix}-{E02}"] })
# VERIFY-MECH (blocked by all feature epics, if selected):
save_issue({ id: "{VERIFY_MECH}", blockedBy: ["{prefix}-{E02}", "{prefix}-{E03}", "{prefix}-{E04}"] })
# HARDEN (blocked by VERIFY-MECH, or all feature epics if no VERIFY-MECH):
save_issue({ id: "{HARDEN}", blockedBy: ["{prefix}-{E02}", "{prefix}-{E03}", "{prefix}-{E04}"] })
# VERIFY-HUMAN (blocked by HARDEN, if selected):
save_issue({ id: "{VERIFY_HUMAN}", blockedBy: ["{prefix}-{HARDEN}"] })
Cross-Milestone Patterns
When one milestone depends on another milestone's completion (e.g. DAT004E depends on DAT004D):
# The VERIFY epic of the dependent milestone is blocked by the HARDEN epic
# of the prerequisite milestone. This is the official cross-milestone dependency wire.
# Example: DAT004E-VERIFY blocked by DAT004D-HARDEN
save_issue({
id: "{DAT004E_VERIFY}",
blockedBy: ["{DAT004D_HARDEN}"]
})
# TRK001A-VERIFY blocked by DAT004D-HARDEN (depends on migrations, not views):
save_issue({
id: "{TRK001A_VERIFY}",
blockedBy: ["{DAT004D_HARDEN}"]
})
Rule: Always wire VERIFY of the downstream milestone to HARDEN of the prerequisite. Never wire individual feature epics to epics in another milestone (too granular, breaks when plans change).
Task-Level blockedBy
Within an epic, tasks that must run sequentially:
# Task 2 blocked by Task 1 (schema must exist before API can reference it):
save_issue({ id: "{task2}", blockedBy: ["{prefix}-{task1}"] })
# REVIEW exit gate task is ALWAYS blocked by ALL other tasks in the epic:
save_issue({
id: "{REVIEW_task}",
blockedBy: ["{t1}", "{prefix}-{t2}", "{prefix}-{t3}"]
})
Tasks that CAN run in parallel have no blockedBy relationship β omit or set blockedBy: [].
3c: Epic Template
For each epic, produce:
## [{milestone_id}] {Title}
**Spec Section:** {spec_file}#Β§{section}
**Dependencies:** {prefix}-{prev}, {prefix}-{prev2}
**Entry Criteria:** {what must be true before starting}
**Exit Criteria:** {what must be true to close}
### Acceptance Tests (written BEFORE tasks)
- [ ] {Testable criterion 1}
- [ ] {Testable criterion 2}
### Tasks
1. {task title} β {size S/M/L} β {labels}
2. {task title} β {size S/M/L} β {labels}
...
### Task Dependency Order
{which tasks can run in parallel vs must be sequential}
3d: Add Verification Epic (FIRST)
ALWAYS add as the FIRST epic in any milestone:
## [{milestone_id}] VERIFY: Prior Milestone Verification
**Dependencies:** None (runs first)
**Entry Criteria:** Previous milestone closed OR first milestone (baseline audit)
### Tasks
1. **Smoke test infrastructure** β verify build, CI, daemon, and core systems are healthy:
- `npm run build` in ui/dashboard/ exits 0
- `pytest tests/` core tests pass
- Factory daemon starts and claims 1 test issue
- Supabase connectivity verified
Size: M β type:test
2. **Quality sweep** β run Phase 6.5 quality sweep on ALL milestone artifacts:
- Milestone issue has required sections
- All epics have Outcome/Tasks/Exit Criteria/Gate Level
- All tasks have Goal/Steps/Outcome/Guardrails/factory labels
- Cross-references intact (no orphans)
Fix anything that fails. Size: M β type:audit
3. Run prior milestone acceptance criteria (regression check) β M β type:test
4. {agents.completion_audit} audit: verify prior milestone completions still hold β M β type:audit
5. {agents.spec_audit} audit: verify spec sections from prior milestone β M β type:audit
6. Fix any regressions found β variable β type:fix
7. Generate baseline test snapshot for this milestone β S β type:test
### Exit Criteria
Infrastructure healthy. All artifacts pass quality sweep. Prior milestone criteria still pass. No regressions. Baseline captured.
For the FIRST-EVER milestone (no prior), this becomes a reality reconciliation:
## [{milestone_id}] VERIFY: Baseline Reality Audit
### Tasks
1. {agents.completion_audit}: audit all items marked DONE in the task-manager milestone β L β type:audit
2. {agents.spec_audit}: spec compliance sweep across all specs β L β type:audit
3. Fix critical lies/gaps found β variable β type:fix
4. Establish baseline acceptance test suite β M β type:test
3d.5: Gate Epic Selection β VERIFY-MECH / HARDEN / VERIFY-HUMAN
Not every milestone needs all three gate epics. Analyze the milestone scope and RECOMMEND
which gates apply. Present to Operator for approval via AskUserQuestion.
Decision logic:
| Milestone touches... | VERIFY-MECH | HARDEN | VERIFY-HUMAN |
|---|---|---|---|
| UI pages, dashboard, Next.js components | β (full 30 modules) | β | β (full 30 flows) |
| Admin CRUD, forms, user-facing features | β (full) | β | β (full) |
| API routes only (no UI) | β (api-contracts, security modules only) | β | β skip |
| Data pipeline, ETL, dbt models | β skip | β (completion + spec audits) | β skip |
| Schema migrations, RLS policies | β (security + database-qa modules only) | β | β skip |
| Infrastructure, CI/CD, monitoring | β skip | β (completion audit only) | β skip |
| Security hardening | β (security modules only) | β | β skip |
| Mobile app | β (cross-browser + responsive) | β | β (mobile-focused flows) |
Ask the Operator:
GATE EPIC RECOMMENDATION for {milestone_id}
ββββββββββββββββββββββββββββββββββββββββββ
This milestone touches: {UI pages / data pipeline / security / etc.}
Recommended gates:
β VERIFY-MECH β {reason: "UI pages need Playwright testing across all roles"}
β HARDEN β {reason: "Always included β completion + spec audits"}
β VERIFY-HUMAN β {reason: "User-facing features need Operator walkthrough"}
or for a data milestone:
β VERIFY-MECH β skip (no UI pages)
β HARDEN β completion + spec audits
β VERIFY-HUMAN β skip (no user-facing changes)
Use AskUserQuestion:
question: "Gate epics for {milestone_id}?"
options:
- "All three (Recommended)" β VERIFY-MECH + HARDEN + VERIFY-HUMAN
- "HARDEN + VERIFY-MECH only" β skip human walkthrough
- "HARDEN only" β audits only, no Playwright
- "Custom" β let me pick
If Operator picks "Custom", ask which modules/flows to include.
If CLAUDE_AUTO=1: use the recommendation without asking.
3e: Add VERIFY-MECH Epic (if selected in 3d.5)
Only add if Operator approved VERIFY-MECH in Step 3d.5. Runs /verify mechanical.
3f: Add Hardening Epic (ALWAYS included)
ALWAYS add after all feature epics. HARDEN runs the completion + spec audits ({agents.completion_audit} + {agents.spec_audit}). If VERIFY-MECH
exists, HARDEN depends on it. If not, HARDEN depends directly on feature epics.
HARDEN runs automated verification via /verify mechanical ONLY if no separate VERIFY-MECH epic exists.
CRITICAL: Milestone-specific outcome criteria. After generating all feature epics,
extract their acceptance tests and translate them into HARDEN-specific test criteria.
These go in a ## Milestone-Specific Test Criteria section in the HARDEN description.
The /verify mechanical skill reads these at runtime to generate targeted Playwright assertions.
# For each feature epic just created:
# Read its ## Acceptance Tests section
# Translate each criterion into a machine-testable assertion:
# "Player can RSVP in/out/maybe" β "POST /api/rsvp returns 200, fact_rsvp row created"
# "Standings update after game" β "After game submit, /standings shows updated W-L record"
# Add to HARDEN description under ## Milestone-Specific Test Criteria
Format in the HARDEN epic description:
## Milestone-Specific Test Criteria (from feature epic acceptance tests)
### From E01: {epic_title}
- [ ] {machine-testable criterion derived from E01's acceptance test 1}
- [ ] {machine-testable criterion derived from E01's acceptance test 2}
### From E02: {epic_title}
- [ ] {criterion}
- [ ] {criterion}
### Cross-Epic Integration
- [ ] {criterion that spans multiple epics β e.g. "player created in E01 appears in standings from E03"}
## [{milestone_id}] HARDEN: Verification & Hardening
**Dependencies:** ALL previous feature epics in this milestone
**Entry Criteria:** All feature epics marked complete
**Skill:** `/verify mechanical --milestone {milestone_id} --severity-gate p0`
**Tool spec:** `.claude/skills/verify/SKILL.md`
**Specs under test:** {list every spec file referenced by this milestone's feature epics}
**Decisions:** {list relevant accepted decisions from storage `list_decisions()` that affect this milestone}
## What This Epic Does
Runs the full `/verify mechanical` automated test suite via Playwright across all roles (commissioner, coach, parent, scorekeeper, anonymous). Uses the existing HARDEN gate epic (created by /plan-milestone) β does NOT create a duplicate. Reads milestone-specific outcome criteria from this description to generate targeted Playwright assertions. Creates child issues for every failure. Screenshots uploaded to Linear (never the repo), console error captures, and step-by-step reproduction logs as issue comments.
## Automated Test Modules (30 total)
### Core Mechanical (7)
- **role-matrix** β every page Γ 5 roles. Expected access vs actual (200/302/403)
- **crud** β create, read, update, delete per admin entity. DB verification after each op. Test data seeded and cleaned up.
- **org-compat** β `?org=`, no `?org=`, invalid `?org=`, subdomain, old `[org]/*` backward compat
- **links** β recursive crawler from `/` and `/demo`, 4 levels deep, 500 page cap. Every `<a>` and `<Link>` followed.
- **data-edges** β empty org, single record, missing fields, stale data
- **responsive** β desktop (1280Γ800), tablet (768Γ1024), mobile (375Γ812) for top 20 pages
- **interactive** β click every button, dropdown, tab, modal, toggle on every page. Verify open/close/populate.
### Security (5)
- **idor** β cross-org data access via direct URL, API call, query param
- **role-escalation** β coach POSTs to admin endpoints, URL param manipulation
- **input-validation** β XSS, SQL injection, oversized input in every form
- **api-contracts** β every `/api/*` route: no-auth=401, wrong-role=403, bad-body=400, valid=200+correct shape
- **permission-mutation** β change role mid-session, toggle privacy, verify immediate effect
### Specialized (7)
- **privacy-coppa** β under-13 default private, parental consent toggles, data deletion
- **user-lifecycle** β invite flow, join code, registration, waitlist, approval/rejection
- **config-fields** β custom JSONB field lifecycle: create β validate β filter β export β delete
- **uploads** β player photos, team logos, CSV import, PDF export, R2 storage verification
- **storage** β R2 file lifecycle, orphan detection, signed URL expiry
- **realtime** β Supabase subscription fire, live update propagation, cleanup on unmount
- **notifications** β RSVP reminders, registration approvals, announcements β verify DB rows
### Resilience (5)
- **chaos** β kill API mid-request, 3G throttle, concurrent edits, corrupt JSONB, token expiry
- **session** β expired token redirect, concurrent sessions, role change mid-session
- **navigation-state** β back/forward, refresh mid-form, deep links, stale tab
- **form-state** β validation, double submit, special chars, max length, navigate away
- **rate-limiting** β 100 rapid requests, verify throttling
### Performance (3)
- **api-performance** β p50/p95/p99 latency per endpoint, 10 concurrent requests, error response shape
- **page-speed** β TTFB, FCP, DOM loaded, full load, transfer size. Score <50 = P2, <25 = P1
- **cache-freshness** β mutation β verify all views reflect change immediately
### Quality (6)
- **accessibility** β axe-core scan, keyboard-only navigation, screen reader flow, color contrast
- **seo-social** β meta tags, OG images, Twitter cards, structured data, sitemap.xml, robots.txt
- **print** β print stylesheet on scoresheets, report cards, bench cards
- **pwa-offline** β service worker, offline fallback, manifest
- **cross-browser** β Chromium, Firefox, WebKit
- **embed-widget** β `/player/[key]/widget` in iframe, CORS, responsive
### Data (4)
- **database-qa** β schema conformance, RLS coverage, orphans, duplicates, migration chain, view freshness, dead tuples
- **database-health** β slow queries (pg_stat_statements >100ms), missing indexes, connection pool
- **data-integrity** β UI prevents invalid data: negative numbers, future birth dates, duplicate slugs
- **spec-compliance** β read specs + decisions, verify each requirement is implemented
### Observability (2)
- **analytics-verify** β PostHog events fire with correct flat-route paths, no PII, org group set
- **monitoring** β /api/health returns 200, Sentry captures errors, background jobs running
## Issue Severity
| Level | Label | Criteria | Blocks Close? |
|-------|-------|----------|---------------|
| P0 | severity:blocker | Security breach, data loss, crash, CRUD broken | YES |
| P1 | severity:major | Feature broken, wrong data, role access wrong | YES |
| P2 | severity:minor | UI wrong, bad copy, layout issue | NO |
| P3 | severity:nit | Spacing, font weight, nice-to-have | NO |
## Process
1. `/verify mechanical` creates `[VERIFY-MECH]` epic in Linear
2. Each test module runs (parallel agents where safe)
3. Failures β child issues with screenshots (uploaded to Linear, never repo)
4. Every step logged as comment on the issue
5. Test data seeded before, cleaned up after
6. Living docs updated incrementally: `docs/testing/SITEMAP.md`, `COVERAGE.md`, `JOURNEYS.md`
7. Results written via `record_test_run` β storage op (supabase β `{storage_schema_qa}.test_runs`)
8. `--fix` mode: auto-fix P0/P1, re-test, close issues (max 3 attempts per issue)
## Tasks
1. Run `/verify mechanical --milestone {milestone_id}` β L β type:test
2. Fix all P0/P1 issues found by /verify β variable β type:fix
3. {agents.completion_audit} audit of all epic completions β M β type:audit
4. {agents.spec_audit} spec compliance check β M β type:audit
5. Fix gaps found by audits β variable β type:fix
6. Re-run `/verify mechanical --severity-gate p0` β must pass clean β S β type:gate
## Exit Criteria
ALL `/verify mechanical` tests pass with zero P0/P1 issues. Completion + spec audits clean. Living docs updated.
3f: Add VERIFY-HUMAN Epic (SECOND-TO-LAST β always included)
ALWAYS add after HARDEN. This is the interactive Operator walkthrough using /verify human.
The mechanical tests (HARDEN) proved the code works. The human tests prove the PRODUCT works.
CRITICAL: Milestone-specific content. The VERIFY-HUMAN description must include:
- Spec references from the feature epics (the specs being TESTED, not the verify skill)
- Outcome-based walkthrough criteria derived from feature epic acceptance tests
- Milestone-specific flows β not just generic "30 flows" but flows tailored to what this milestone built
# After generating all feature epics, build the VERIFY-HUMAN content:
# 1. Collect all spec refs from feature epics' **Spec Section:** lines
# 2. Collect all acceptance tests from feature epics
# 3. Translate each into a human-walkable outcome:
# Machine: "POST /api/rsvp returns 200"
# Human: "Parent taps 'I'm In' β confirmation shown, RSVP count updates"
# 4. Group by the 30-flow categories (role walkthroughs, CRUD, journeys, etc.)
# 5. Add a ## Milestone-Specific Specs section listing every spec this milestone touches
## [{milestone_id}] VERIFY-HUMAN: Operator Walkthrough
**Dependencies:** [{milestone_id}] HARDEN epic must be Done
**Entry Criteria:** All mechanical tests pass, all audits clean
**Skill:** `/verify human --milestone {milestone_id}`
**Tool spec:** `.claude/skills/verify/SKILL.md`
**Specs under test:** {list every spec file referenced by this milestone's feature epics}
**Decisions:** {list relevant accepted decisions from storage `list_decisions()`}
**Resume:** `/verify human --continue {task_prefix}-{this_epic_id}`
## What This Epic Does
Interactive Playwright walkthrough with the Operator. The agent drives the browser, navigates each page, shows it to the Operator, waits for feedback, fixes issues in real-time (edit β commit β re-test), and marks each flow Done. This is NOT "create issues and leave" β it's a live session.
The HARDEN epic proved the CODE works. This epic proves the PRODUCT works β UX, intuitiveness, flow, feel, design, and data correctness through human eyes.
## How It Works
1. Agent creates child issues for each walkthrough flow under this epic
2. For each flow: agent navigates via Playwright, screenshots each step, describes what it sees
3. Operator watches and gives feedback ("fix that", "try on mobile", "what about as a coach?")
4. Agent fixes bugs immediately β edit, commit, re-test, close the bug issue
5. Every step logged as a comment on the flow's Linear issue
6. Screenshots uploaded to Linear (never the repo)
7. When a flow is complete, agent marks that flow's issue Done
8. Session can pause anytime β resume with `/verify human --continue {task_prefix}-{this_epic_id}`
## 30 Walkthrough Flows (mirrors mechanical modules through human lens)
### Role Walkthroughs (5) β does each role's view make sense?
1. Commissioner: Admin overview β dashboard layout, quick actions, data summary
2. Coach: Coach overview β practice tools, player access, eval access
3. Parent: Parent overview β kid's info, schedule, stats visibility
4. Scorekeeper: Scorekeeper overview β game list, start game flow
5. Public: Anonymous browse β landing page, what's visible without login
### CRUD Usability (6) β are the forms obvious? Fields in right order?
6. Player CRUD β add, edit, delete, search, filter, export
7. Team CRUD β add, edit, roster assignment
8. Game CRUD β add game, edit, cancel
9. Season CRUD β create, configure, open/close
10. Registration CRUD β open, approve, reject, waitlist
11. User CRUD β invite, role change, deactivate
### User Journeys (4) β would a real person complete this without help?
12. Registration: landing β signup β join code β approve β access
13. Game Day: lineup β track β score β stats update
14. Season Setup: create β register β draft β schedule β play
15. Config Change: toggle feature β verify across all roles
### UI/UX/Mobile (5) β does mobile feel native? Is the nav logical?
16. Mobile: Parent at rink β RSVP, schedule, live score on phone (375Γ812)
17. Mobile: Scorekeeper β game tracking on tablet (768Γ1024)
18. Navigation: Can you find it? β 10 tasks, measure clicks to complete
19. Error UX: Trigger 5 error types β are the messages helpful or confusing?
20. First Impression: Fresh eyes β would you pay for this? What's confusing?
### Entity Pages (6) β is the data correct and well-presented?
21. Player profile β stats, tabs, sub-pages, sharing, portfolio
22. Team page β roster, standings context, schedule
23. Game detail β box score, play-by-play, analytics
24. Standings β sort, filter, season selector
25. Leaders β leaderboard accuracy, filtering
26. Schedule β calendar view, upcoming/past
### Config & Customization (4) β can an admin figure it out?
27. Custom fields: create field β appears on form β filters β export
28. Feature toggles: toggle off β hidden for all roles β toggle on
29. Org settings: name, logo, timezone β reflected everywhere
30. Privacy controls: make player private β verify public can't see
## Issue Structure
Each walkthrough flow = a Linear issue under this epic:
[VERIFY-HUMAN] {Role}: {Flow Name} Labels: verify-finding, module:human, role:{role}
Bugs found during walkthrough = separate child issues:
[VERIFY-HUMAN] P{severity} β {page} β {description} Labels: verify-finding, severity:{level}, module:human
Every step logged as a comment:
Step 3/8 β {timestamp}
Action: Click 'Add Player' button Result: Modal opened β "Add New Player" form Screenshot: [attached] Verdict: β PASS
Operator feedback: "fields should be in order: name, position, jersey, team" β Created: BEN-XXXX β P3 β field order on add player modal
## Session Continuity
At break points, the agent prints progress and offers to pause:
VERIFY-HUMAN β {task_prefix}-{epic_id} β 8/30 flows complete Bugs: 5 found | 3 fixed | 2 open Resume: /verify human --continue {task_prefix}-{epic_id}
Any session can resume β Linear tracks what's done and what's left.
## Spec-Informed Testing
Before testing each feature, the agent reads:
- Relevant spec from `{spec_dir}/`
- Accepted decisions from storage `list_decisions()`
- Epic acceptance criteria from the milestone's feature epics
- Tests against ALL of the above β not just "does it render"
## Tasks
1. Run `/verify human --milestone {milestone_id}` β L β type:test
2. Fix all bugs found during walkthrough β variable β type:fix
3. Re-test fixed issues with Operator β S β type:gate
4. Operator sign-off on all 30 flows β S β type:gate
## Exit Criteria
ALL 30 walkthrough flows marked Done in Linear. Operator satisfied with UX, design, and data correctness. All P0/P1 bugs fixed. Living docs updated.
Z Epic β DEPRECATED (removed 2026-05-29)
The Z (E2E Verification) epic is no longer created. Its responsibilities are fully covered by:
- VERIFY-MECH β automated Playwright testing (replaces Machine Proof)
- VERIFY-HUMAN β interactive Operator walkthrough (replaces Human Proof)
The old --z-epic flag and E2E Mode settings are ignored. CCB Phase 3.25 no longer
asks about Z epics.
Step 3g: Assign Machine / Agent / Model per Epic
Every epic gets execution context assigned. This comes from CCB Phase 3.3 decisions
(if called from /ccb) or is auto-assigned using these defaults.
Goal: keep costs down and all machines running. Use the cheapest model/agent that can handle the task. Distribute work across machines to maximize throughput. Only use opus for work that genuinely requires deep reasoning. Prefer free/cheap agents for routine code generation. Ollama is zero-cost and should be the first choice for any task it can handle.
Agent & Model Catalog
Full lifecycle process for all agents:
.agents/reference/lifecycle_process.md
| Agent | Sub-Models | Machine Constraint | Cost | Best For |
|---|---|---|---|---|
| claude | opus, sonnet, haiku | any | $$$/$$/$ | Full tool access, complex reasoning, implementation |
| gemini | 2.5-pro, 2.5-flash, 2.0-flash | any | $$/$/ free | Architecture, specs, review, large context |
| codex | gpt-4.1, gpt-4.1-mini | any (cloud) | free | Standard code gen, tests, bulk work |
| cursor | sonnet, gpt-4.1 | mothership | $$/free | IDE-based UI/React, multi-file edits |
| ollama | qwen3:32b, llama3.3:70b, deepseek-r1:32b, codestral:22b, mistral:7b | mothership ONLY | free | Boilerplate, tests, docs, small refactors, code completion |
Cost tiers:
- Free: codex/gpt-4.1, codex/gpt-4.1-mini, ollama/*, gemini/2.0-flash
- Cheap: claude/haiku, gemini/2.5-flash
- Standard: claude/sonnet, gemini/2.5-pro, cursor/sonnet
- Premium: claude/opus (use sparingly β architecture, audits, gates only)
Ollama sub-model selection:
qwen3:32bβ default, best general code quality at 32Bcodestral:22bβ pure code completion, fastdeepseek-r1:32bβ reasoning-heavy tasks (mini-opus)llama3.3:70bβ largest context, slower, use for bigger tasksmistral:7bβ fastest, use for trivial tasks only
Epic-Level Routing Defaults
Every epic gets Primary + Backup 1 + Backup 2. All three must use DIFFERENT agents.
| Epic Type | Primary | Backup 1 | Backup 2 | Why |
|---|---|---|---|---|
| VERIFY | workhorse/claude/sonnet | mothership/ollama/qwen3:32b | auditor/gemini/2.5-flash | Routine checks β try local first |
| Feature (standard code) | workhorse/codex/gpt-4.1 | mothership/ollama/qwen3:32b | auditor/gemini/2.5-flash | Free tier cascade |
| Feature (boilerplate/CRUD) | mothership/ollama/codestral:22b | workhorse/codex/gpt-4.1 | auditor/gemini/2.0-flash | Cheapest possible |
| Feature (architecture/spec) | workhorse/claude/opus | mothership/gemini/2.5-pro | workhorse/codex/gpt-4.1 | Deep reasoning needed |
| Feature (UI/React) | mothership/cursor/sonnet | workhorse/codex/gpt-4.1 | mothership/ollama/qwen3:32b | Cursor excels at UI |
| Feature (CV/ML) | mothership/ollama/deepseek-r1:32b | workhorse/claude/sonnet | mothership/gemini/2.5-pro | Local inference + reasoning |
| Feature (data/ETL) | workhorse/codex/gpt-4.1 | mothership/ollama/qwen3:32b | auditor/gemini/2.5-flash | ETL is standard code |
| Feature (Python/complex) | workhorse/claude/sonnet | mothership/ollama/qwen3:32b | auditor/gemini/2.5-flash | Needs Python expertise |
| HARDEN (audits) | workhorse/claude/opus | mothership/gemini/2.5-pro | workhorse/codex/gpt-4.1 | Deep reasoning |
| Z (E2E) | mothership/claude/opus | workhorse/gemini/2.5-pro | mothership/cursor/sonnet | Operator-interactive |
Backup rules:
- Primary, Backup 1, and Backup 2 must ALL use different agents (never claudeβclaudeβgemini)
- Backup machine SHOULD differ from primary when possible
- Backup model must be from its agent's ecosystem
- If primary fails/is unavailable, factory falls back to Backup 1 automatically, then Backup 2
Ollama Emergency Fallback (MANDATORY):
Every epic and every task MUST specify an Ollama model in its Execution Context,
even if Ollama isn't Backup 1 or 2. This is the "tokens ran out" safety net.
Format: Ollama fallback: mothership/ollama/{model} β always mothership, always free.
- Default:
mothership/ollama/qwen3:32b(best general quality) - For pure code tasks:
mothership/ollama/codestral:22b - For reasoning-heavy:
mothership/ollama/deepseek-r1:32b - For large context:
mothership/ollama/llama3.3:70b
Milestone-Level Backup
In addition to per-epic assignments, assign ONE milestone-level backup β the agent+model you'd use if you had to run the ENTIRE milestone on a single fallback:
## Milestone Backup
If Claude is unavailable for this milestone:
Fallback: {machine}/{agent}/{model} (e.g. workhorse/codex/gpt-4.1)
Ollama tasks: mothership/ollama/qwen3:32b (always available, zero cost)
This goes in the milestone description in Linear.
Task-Level Routing Overrides
Tasks inherit epic defaults but can override. Route by task characteristics:
| Task Type | Model Override | Why |
|---|---|---|
| size:S + type:test | ollama/codestral:22b or haiku | Simple test scaffolds β try local first |
| size:S + type:docs | ollama/mistral:7b or gemini/2.0-flash | Cheapest for docs |
| size:S + non-critical | ollama/qwen3:32b or codex/gpt-4.1-mini | Zero/low cost |
| size:M + standard code | codex/gpt-4.1 or ollama/qwen3:32b | Cost-effective code gen |
| size:M + Python-heavy | claude/sonnet or ollama/qwen3:32b | Needs Python expertise |
| size:L + architecture | claude/opus | Worth the cost |
| type:audit | claude/opus or gemini/2.5-pro | Must be thorough |
| type:gate | claude/opus | Critical decision point |
| Schema/migration | claude/opus | Correctness critical |
Machine Distribution
Spread work across machines to maximize parallelism:
- Mothership: Operator-interactive tasks, ALL Ollama tasks, Cursor tasks, CV pipeline
- Workhorse: Always-on factory tasks (claude/codex), bulk code gen, tests, ETL
- Auditor: Review tasks, gemini tasks, lightweight builds
Ollama runs ONLY on mothership β it requires the local GPU. Never assign ollama to workhorse or auditor.
Never pile all tasks on one machine when others are idle.
Add to epic labels: agent:{agent}, model:{model}, machine:{machine}
These propagate to child tasks as defaults (tasks can override).
Operator can override any assignment during Step 4 review.
Step 4: Operator Review (Interactive)
Present the full plan to Operator:
Milestone {milestone_id}: {goal}
Acceptance Criteria: {count}
Epics: {count} ({count} feature + 1 hardening)
Total Tasks: {count}
Estimated Factory Time: {estimate}
Epic Order:
1. [{milestone_id}] VERIFY ({task_count} tasks)
2. [{milestone_id}] {title} ({task_count} tasks, {deps})
3. [{milestone_id}] {title} ({task_count} tasks, {deps})
...
N. [{milestone_id}] HARDEN ({task_count} tasks)
Ask Operator:
- "Does this epic ordering make sense?"
- "Any tasks missing or unnecessary?"
- "Should any epics be split or merged?"
- "Priority labels for the tasks?" (P0/P1/P2)
Step 5: Create Artifacts (Linear Only)
Linear is the sole planning/tracking layer. No GitHub milestones or tracking issues. Epics and tasks go to Linear. The factory webhook reads from the task manager queue. GitHub is for code, PRs, and CI only.
BATCH STRATEGY β MANDATORY to maintain quality across all tasks: Do NOT write task descriptions on-the-fly during
save_issuecalls. That pattern causes later tasks to be thin because context is exhausted. Instead:
- Draft ALL task descriptions in conversation first (before any
save_issuecalls). Write each task body in full β all 15+ sections with real content.- Review the full set for quality and consistency.
- Then create in Linear β at that point you're copy-pasting complete descriptions, not generating under pressure.
This ensures task #1 and task #30 receive equal attention.
5a-pre: Resolve Project & Milestone in Linear (CRITICAL)
Before creating ANY issue, resolve the Linear project and milestone IDs.
# 1. Find the Linear project by code prefix
projects = list_projects(team: "{task_team_id}")
# Match project by milestone code prefix (e.g., CVX β "Computer Vision", ETL β "Pipeline / ETL")
# Store: {project_id}
# 2. Find or create the Linear milestone
milestones = list_milestones(team: "{task_team_id}")
# Match milestone by name containing {milestone_id} (e.g., "CVX005E")
# If not found, create it:
# save_milestone(name: "{milestone_id}: {goal}", team: "{task_team_id}")
# Store: {milestone_id_linear}
CONFIRMATION GATE β show the user before proceeding:
LINEAR TARGET CONFIRMATION
βββββββββββββββββββββββββ
Team: {task_team_id}
Project: {project_name} ({project_id})
Milestone: {milestone_name} ({milestone_id_linear})
Labels: milestone:{milestone_id}, epic, type:{type}
Creating: {epic_count} epics, {task_count} tasks
Confirm? (y/n)
Use AskUserQuestion to confirm. Do NOT create any issues until confirmed.
If the project or milestone looks wrong, let the user correct it before proceeding.
5a: Create Epic Issues in Linear
DEDUP CHECK: Before creating ANY epic, search Linear:
list_issues(team: "{task_team_id}", query: "E{number}")
If a matching epic exists, update it with save_issue(id: "{n}",...).
For each epic, call save_issue with:
title: "[{milestone_id}] {title}"
team: "{task_team_id}"
state: "Backlog"
priority: {1=Urgent, 2=High, 3=Normal, 4=Low}
projectId: {project_id} β MUST be set
milestoneId: {milestone_id_linear} β MUST be set
labels: ["epic", "priority:{level}", "machine:{machine}", "type:{type}", "milestone:{milestone_id}"]
description: (markdown β same structure as before, see below)
blockedBy wiring β do NOT skip: After all epics are created, wire
blockedByrelations per Section 3b.1. See Step 5a.5 below for the two-pass algorithm. The Linear UI "Blocked by" widget,/epic startdependency check, andlinear_epic_ops.py nextall read the structured field β markdown bullets alone are invisible to these tools.
Epic description body (markdown β ALL placeholders must be replaced with real content, minimum 2-3 sentences per section):
## Goal
{2-4 sentences: WHAT this epic delivers, WHY it matters for the milestone, and what breaks if it's skipped}
## Milestone
{milestone_id}: {goal}
## Git
- **Branch:** `{type}/{prefix}-{epic_id}-{short-kebab-name}`
- **PR target:** `develop`
- **PR scope:** All tasks in this epic ship as one PR
- **Commit prefix:** `[{TYPE}] {prefix}-{task_id}: {what changed}`
## Outcome
- **Measured By:** {SQL query, test command, or metric that proves completion}
- **Baseline:** {current state}
- **Target:** {what "done" looks like in numbers}
## What Should Happen When This Epic Is Done
- {Concrete outcome 1}
- {Concrete outcome 2}
**Spec Section:** {spec_file}#Β§{section}
**Dependencies:** {deps β list {prefix}-{n} identifiers}
### Acceptance Tests
`tests/milestones/m{id}/test_e{number}_{name}.py`
- [ ] {test 1}
- [ ] {test 2}
### Tasks
_Task {prefix}-{n} identifiers filled after task creation._
### Exit Criteria
{criteria}
---
## Execution Context
- **Machine:** {machine}
- **Agent:** {agent}
- **Model:** {model}
- **Backup 1:** {backup1_machine}/{backup1_agent}/{backup1_model}
- **Backup 2:** {backup2_machine}/{backup2_agent}/{backup2_model}
- **Ollama fallback:** mothership/ollama/{ollama_model}
## Process Reference
All agents MUST read `.agents/reference/lifecycle_process.md` before starting work.
This file contains the full branch, commit, PR, and review process β agent-agnostic.
## Lifecycle Instructions (for all agents β Claude, Gemini, Codex, Cursor, Ollama)
### Starting This Epic
1. Branch from `develop`: `git checkout -b {branch_name}`
2. All tasks are commits on this branch β do NOT create sub-branches
3. Commit format: `[{TYPE}] {prefix}-{task_id}: {description}`
4. If using Claude Code: run `/epic start {prefix}-{epic_id}`
### Working Tasks
1. Read the task's ## Required Reading section first
2. Follow ## Steps exactly
3. Run the test in ## Outcome after each task
4. Commit immediately after each task passes
5. Mark each task Done in Linear when its commit is pushed
### Closing This Epic
1. Run acceptance tests: `pytest {test_file} -v`
2. Create PR to `develop` with title: `[{MILESTONE}] {Epic Title}`
3. Wait for CodeRabbit + assigned model reviewers
4. Fix any blocking review feedback, push, re-request review
5. Do NOT merge β wait for Operator approval (unless autonomy:green)
6. After merge: mark all tasks and this epic as Done in Linear
7. Post "What Actually Happened" comment on this epic: delivered items, PR link, test results, deferred items
8. **Start a new session for the next epic**
9. If using Claude Code: run `/epic close {prefix}-{epic_id}`
---
## What Actually Happened
_To be completed at epic close._
## Auditor Comments
_To be completed at epic close._
Record the returned id (e.g. {prefix}-7) β you need it as parentId for tasks.
5a-post: Spec Coverage Producer Hook (REQUIRED β D-1427)
Immediately after each successful save_issue epic
creation, write the epic's spec-coverage row to spec-coverage storage (supabase β {storage_schema_qa}.spec_section_coverage).
This is the producer half of the Β§46 spec-coverage loop. Locked by
D-1427.
For each epic created in 5a, build a JSON payload from the epic and pass it on stdin to Python β never embed the epic description as a raw triple-quoted string (epic bodies frequently contain quotes, backticks, and shell metacharacters that would break the literal):
EPIC_PAYLOAD=$(python3 -c "import json; print(json.dumps({
'epic_id': '{epic_id}',
'epic_status': 'Backlog',
'epic_description': '''{epic_description_python_safe}''',
}))")
echo "$EPIC_PAYLOAD" | python3 -c "
import sys, json
from scripts.factory.spec_coverage_writer import upsert_planned_coverage_from_description
payload = json.load(sys.stdin)
results = upsert_planned_coverage_from_description(**payload)
for r in results:
print(json.dumps(r))
"
If the calling agent already has the epic body in a Python variable
(epic["description"]), prefer calling
upsert_planned_coverage_from_description(...) directly β the
JSON-via-stdin shape exists only for shell-based invocations where the
body needs to cross a shell boundary safely.
The hook is non-blocking β if Supabase is unreachable or the
section row is missing from the scanner ledger, the writer logs a
warning and returns a result dict; it never raises. Epic creation
proceeds regardless. Skipped silently when the epic body has no
**Spec Section:** line (epics without spec refs are valid).
Print one line per (spec_file, section_id) pair: coverage: {epic_id} {spec_file}#{section_id} β ok|warning|error. After all epics, expect
β₯1 ok per epic that carries a **Spec Section:** line.
5a.5: Wire Epic blockedBy Relations (Two-Pass β REQUIRED)
Epic blockedBy cannot always be set on create because a blocker may not exist yet (forward
reference). Use this two-pass strategy after ALL epics are created:
Pass 1 β already done in Step 5a: create every epic, collect returned identifiers in a map:
epic_ids = {
"VERIFY": "{ID}",
"E01": "{task_prefix}-{n+1}",
"E02": "{task_prefix}-{n+2}",
...
}
Pass 2 β wire blockedBy for each epic that has planned blockers:
# Example: 3-epic chain E1 β E2 β E3, all gated by VERIFY, HARDEN at the end
save_issue({ id: epic_ids["E01"], blockedBy: [epic_ids["VERIFY"]] })
save_issue({ id: epic_ids["E02"], blockedBy: [epic_ids["VERIFY"]] })
save_issue({ id: epic_ids["HARDEN"], blockedBy: [epic_ids["E01"], epic_ids["E02"]] })
save_issue({ id: epic_ids["Z"], blockedBy: [epic_ids["HARDEN"]] })
Rules:
blockedByis append-only β callingsave_issuewithblockedByon an existing epic ADDS relations without removing existing ones. Safe to call multiple times.blocks(inverse) is auto-populated by Linear β never set it explicitly.- Markdown
## Blocked bysections STILL appear in epic descriptions β relations supplement, they do not replace the human-readable markdown. Both must agree. - Cross-milestone deps follow Section 3b.1 pattern: wire the dependent milestone's VERIFY epic to the prerequisite milestone's HARDEN epic.
Verification after Pass 2: For each epic with blockers, call
get_issue(id: "{n}", includeRelations: true)and confirmrelations.blockedByis non-empty. If empty, the wire was silently skipped β retry.
5d: Create Task Issues in Linear (DoR-Compliant)
QUALITY WARNING β READ BEFORE CREATING ANY TASK: Every task description MUST be substantial. The factory agent reads ONLY the Linear issue β there is no other source of truth. Thin descriptions = broken factory runs.
Minimum content requirements (enforced β not advisory):
## Goalβ 2-4 sentences explaining WHAT and WHY (not just the title restated)## Contextβ 3+ sentences: parent epic relationship, spec motivation, what depends on this## Stepsβ 3-8 numbered steps, each with a specific file path or function name. Every step must be atomic (one green commit). Never write "implement X" without naming the file.## Required Readingβ at minimum 3 real file paths that exist in the repo (verify withtest -f)## Pre-Answered Questionsβ at minimum 2 Q&A pairs that pre-answer what a factory agent would ask## Acceptance Criteriaβ at minimum 3 checkboxes with testable conditions## Spec Updatesβ either list >=1 spec file + section, or write "None β no spec impact" with justification## Testsβ use the layer routing table (Step 2b) to identify which layer(s) apply to what THIS task builds. For every layer that applies, the test file is MANDATORY β not optional, not deferred, not "TODO." The file must exist on disk and fail (RED) before this task is created. Writing "None" is only valid if no layer in the routing table matches β justify it.Anti-patterns that will FAIL review (do not do these):
- Placeholder text like "{one sentence}" or "{why}" left unfilled
- Steps that say "implement X" without naming files or functions
## Required Readingwith fewer than 3 files## Stepswith fewer than 3 numbered items## Goalthat just restates the title- Any section left at its template default
You are writing for a factory agent that has ZERO context beyond this ticket. Write as if you are handing this to a smart developer starting on day 1 with no prior knowledge. Every ambiguity you skip = a wrong implementation or a stuck agent.
DEDUP CHECK: Search Linear for existing tasks before creating:
list_issues(team: "{task_team_id}", query: "{task title keywords}")
For each task, call save_issue with:
title: "{title}"
team: "{task_team_id}"
parentId: "{epic_identifier}" β e.g. "{{prefix}}-7" (the epic's Linear ID)
projectId: {project_id} β MUST match the epic's project
milestoneId: {milestone_id_linear} β MUST match the epic's milestone
state: "Todo"
priority: {1-4}
labels: ["task", "model:{model}", "machine:{machine}", "priority:{level}", "autonomy:{level}", "milestone:{milestone_id}"]
description: (full DoR markdown body β see below)
Task description body (markdown β ALL sections required, ALL placeholders must be filled with real content):
## Goal
{One sentence: what does this task achieve and why}
## Context
{Why this matters β link to parent epic, what depends on this, spec motivation}
Parent epic: {prefix}-{epic} β {epic_title}
{If blocked by prior task: "Blocked by {prefix}-{prev} which delivers {what}"}
## Git
- **Branch:** Work on parent epic branch `{type}/{prefix}-{epic_id}-{short-kebab-name}` β do NOT create a separate branch
- **Commit prefix:** `[{TYPE}] {prefix}-{task_id}: {what changed}`
## Spec Ref
{spec_file}#{section}
## Required Reading
- `{file_path_1}` β {why}
- `{file_path_2}` β {why}
- `rules/areas/{area}.md` β area-specific rules
## Pre-Answered Questions
- Q: {question} β A: {answer}
## Steps
1. Read Required Reading files above
2. {Concrete step with file path and function name}
3. {Next step}
4. {Verification: "Run `{test_command}` β expect {outcome}"}
> **Step authoring rule:** Each step must be atomic enough to be a single green commit.
> After each step: run its test, then commit. Max 5 files per step.
> If a step would touch >5 files or take >2 hours, split it into sub-steps.
> If a step needs a database, note: "Start `sandbox-postgres` before this step."
## Acceptance Criteria
- [ ] {testable criterion 1}
- [ ] {testable criterion 2}
## Outcome
- **Measured By:** {test command or file check}
## Guardrails
- {Scope limits}
- Do not modify files outside the scope of this task
## Spec Updates
Update these spec files if your changes affect their documented behavior:
- `{spec_file_1}` β Β§{section}: {what to update}
- `{spec_file_2}` β Β§{section}: {what to update}
_Spec updates MUST be included in the same PR as the code change. Do not merge code that makes a spec inaccurate._
## Tests
- Layer {N} ({layer_name}) β `tests/{suite}/test_{milestone_id_lower}_{epic_short}_{what}.py`
- Asserts: {exact outcome this test checks β derived from exit criteria above}
- Run: `pytest {test_file_path} -x --tb=short`
_Only include layers that match what this task actually builds (see layer routing table in Step 2b).
A SQL migration doesn't get a Playwright test. A hook doesn't get a canonical golden test.
Pick the 1-2 layers that directly prove this specific task's exit criterion β no more.
File must already exist on disk, written RED (failing) before tasks execute β see Step 2b.
completion-audit reads this section and runs these files before accepting Done/Complete.
Pre-commit guard blocks Done commits if the file is missing or pytest fails._
## Agents to Call
- {agent_1} β {what to review}
- code-reviewer β final code quality check
## TDD
- **Test file:** `tests/milestones/{milestone_id}/test_{epic_short_name}.py`
- **Test class:** `Test{MilestoneId}{EpicName}`
## Execution Context
- **Model:** {model}
- **Machine:** {machine}
- **Agent:** {agent}
- **Backup 1:** {backup1_machine}/{backup1_agent}/{backup1_model}
- **Backup 2:** {backup2_machine}/{backup2_agent}/{backup2_model}
- **Ollama fallback:** mothership/ollama/{ollama_model}
## Process Reference
**READ FIRST:** `.agents/reference/lifecycle_process.md` β full agent-agnostic process.
**Area rules:** `rules/areas/{area}.md`
**Core rules:** `rules/BASE.md`
## Lifecycle Instructions (for all agents β Claude, Gemini, Codex, Cursor, Ollama)
1. Work on parent epic branch `{branch_name}` β do NOT create a new branch
2. Commit format: `[{TYPE}] {prefix}-{task_id}: {description}`
3. **Commit after EACH step β not just at task completion.** Run the step's test first, then commit. Max 5 files per commit.
4. If a step requires a database, start `sandbox-postgres` first: `docker compose up -d sandbox-postgres`
5. Run `{test_command}` after implementation β must pass
6. Mark this task Done in Linear when complete
7. If using Claude Code and this is the last task: run `/epic close {prefix}-{epic_id}`
---
**Size:** {S|M|L}
**Blocked by:** {{prefix}-{prev} or "None"}
### Files to Modify
- `{file_path}` β {what changes}
### Files to Create
- `{file_path}` β {purpose}
5d.0: Project-Specific Agent Routing (domain-specific β override in project overlay skill)
Use these when task domain matches. Each entry includes the agent definition file path so non-Claude models (Gemini, Codex) can read the agent spec directly.
| Task Domain | Agent to Call | Agent File (absolute path) |
|---|---|---|
| Supabase migrations, RLS policies, Realtime config | supabase-specialist | .claude/agents/07-specialized-domains/supabase-specialist.md |
| dbt models, mart views, ETL transforms, stageβfact | etl-specialist | .claude/agents/07-specialized-domains/etl-specialist.md |
| FastAPI routes, Pydantic models, API endpoints | backend-developer | .claude/agents/01-core-development/backend-developer.md |
| Tracker / Scorekeeper UI (v30 event model, React) | tracker-specialist | .claude/agents/07-specialized-domains/tracker-specialist.md |
| Dashboard pages, live feed, Next.js pages | dashboard-developer | .claude/agents/07-specialized-domains/dashboard-developer.md |
| Playwright E2E tests, UI smoke tests | ui-comprehensive-tester | .claude/agents/07-specialized-domains/ui-comprehensive-tester.md |
| Completion reality audit (is it actually done?) | {agents.completion_audit} (default: completion-audit) | .claude/agents/07-specialized-domains/completion-audit.md |
| Spec compliance audit (does code match spec?) | {agents.spec_audit} (default: spec-audit) | .claude/skills/spec-audit/SKILL.md |
| Code quality review, PR review | code-reviewer | .claude/agents/04-quality-security/code-reviewer.md |
| IndexedDB, offline/sync, frontend state | frontend-developer | .claude/agents/01-core-development/frontend-developer.md |
| Hockey domain logic, event types, stat rules | hockey-analytics-sme | .claude/agents/07-specialized-domains/hockey-analytics-sme.md |
| Computer vision pipeline, XY tracking | cv-engineer | .claude/agents/05-data-ai/cv-engineer.md |
| Python ETL, pandas, calculations | python-pro | .claude/agents/02-language-specialists/python-pro.md |
| TypeScript, React, Next.js components | typescript-pro | .claude/agents/02-language-specialists/typescript-pro.md |
How to use in task descriptions:
The ## Agents to Call section in each task body MUST list the specific agent names from this table
(not just "reviewer" or "audit agent"). Format:
## Agents to Call
- `supabase-specialist` β review migration SQL for correctness and RLS policy
File: `.claude/agents/07-specialized-domains/supabase-specialist.md`
- `code-reviewer` β final code quality check on all changed files
File: `.claude/agents/04-quality-security/code-reviewer.md`
Non-Claude models: to read an agent's capabilities before calling it, Read the file at the
path listed above. All agent files are relative to the repo root:
{project_root}/
5d.1: Auto-Populate DoR Sections
Before generating each task body, resolve these sections from plan context:
Required Reading β auto-discover from:
- The task's area β map to key files using
areasfrom.claude/project-context.md. If the project defines areaβpath mappings, use them. Otherwise, usegrep -rl {area_keyword}.to find relevant files.Project overlay: define your area mappings in
.claude/skills/plan-milestone/SKILL.mdso this step uses project-specific paths instead of generic discovery. - The spec ref β include the spec file itself
grep -rlfor function/class names mentioned in the task steps β include those files- Relevant decisions from your project's decision log (location defined in CLAUDE.md)
Every file listed MUST exist in the repo. Run test -f {path} before including.
Model β route by complexity (cheapest that works):
size:S+type:testβhaiku(cheapest)size:S+type:docsβflash(cheapest)size:S+ non-critical βsonnetsize:M+ standard code βsonnetorgpt-4.1(if epic uses codex)size:Lor architecture/spec-heavy βopustype:auditβopustype:gateβopus- Schema/migration work β
opus
Agent β inherit from epic, override if task needs differ:
- Tasks inherit parent epic's agent by default
- Override only when task requires a different capability
- E.g.: UI task in a data epic β cursor override
Machine β distribute across fleet:
- Default: inherit from parent epic
- Docker/container tasks β
workhorse - Tasks needing Ollama β
mothership - Interactive/Operator tasks β
mothership - Lightweight review/test tasks β
auditor(keep it busy) - Never stack all tasks on one machine
5d.2: DoR Self-Check (MANDATORY before creating)
Before creating each task, verify ALL of these are true. Do not skip β this is the gate.
DoR CHECKLIST (all must pass β quality, not just presence):
[ ] Goal: 2-4 sentences, explains WHAT and WHY β not just the title restated
[ ] Context: 3+ sentences referencing parent epic, spec motivation, and what depends on this
[ ] Spec Ref: a real spec file path (test -f passes) β not a placeholder
[ ] Required Reading: >=3 files that exist in the repo (verified with test -f)
[ ] Pre-Answered Questions: >=2 Q&A pairs answering what a factory agent would ask
[ ] Steps: >=3 numbered items, each naming a specific file path or function
[ ] Steps: each step is atomic enough to be one green commit (max 5 files)
[ ] Acceptance Criteria: >=3 checkboxes with testable, specific conditions
[ ] Outcome: a concrete verification command (e.g., "pytest tests/... expect X passing")
[ ] Guardrails: >=1 specific scope limit (files NOT to touch, behaviors NOT to add)
[ ] Agents to Call: >=1 agent from the project routing table with file path
[ ] TDD: names a specific test file path and class name (not "TBD" or generic)
[ ] Model: set to a real model (not inherited placeholder)
[ ] Machine: set to a specific machine (mothership/workhorse/auditor)
[ ] Spec Updates: real file+section listed, or "None β no spec impact" with justification
[ ] Tests: layer number + file path matching what THIS task builds (use routing table) β file exists on disk and fails (RED)
[ ] Labels: task, model:*, machine:*, priority:* all present
[ ] NO unfilled placeholders: scan for {curly_braces} β all must be replaced with real values
If ANY item fails: rewrite the failing sections before calling save_issue. Do not create a thin ticket and plan to fix it later β it will fail the CCB quality sweep and require a re-run.
5e: Create Epic Exit Gate Task (MANDATORY for every epic)
Every epic MUST have a final review task as a child of the epic in Linear:
save_issue({
title: "{prefix}-{epic}-REVIEW Exit gate verification",
team: "{task_team_id}",
parentId: "{epic_identifier}",
state: "Backlog",
priority: 2,
labels: ["task", "model:opus", "machine:any"],
description: "## Goal\nVerify epic {prefix}-{epic} achieved what it claimed...\n\n## Steps\n1. Run acceptance tests\n2. {agents.completion_audit} audit\n3. Fill What Actually Happened\n..."
})
The HARDEN epic does NOT get an exit gate task β it IS the exit gate. The VERIFY epic DOES get one.
5e.1: Specialist Domain Review (MANDATORY for every epic)
After all tasks for an epic are drafted but BEFORE calling save_issue for any of them,
spawn the relevant domain specialist(s) to review the task set. This catches scope gaps,
wrong file targets, and missing guardrails that the completion/spec audits don't cover.
Routing table β which specialist for which epic content:
| Epic touches... | Specialist to call |
|---|---|
| Any SQL migration, RLS, Supabase schema | supabase-specialist |
ETL pipeline, pandas, src/calculations/, dbt | etl-specialist |
Next.js dashboard, React components, ui/dashboard/ | dashboard-developer |
| Tracker UI, v30 event model, scorekeeper | tracker-specialist |
CI/CD, GitHub Actions, .github/workflows/ | devops-engineer |
| Hockey stat logic, xG, Corsi, Fenwick | hockey-analytics-sme |
CV pipeline, src/cv/, camera calibration | cv-engineer |
| Python code quality, modularization | python-pro |
| TypeScript, strict mode, component patterns | typescript-pro |
Always call regardless of domain:
{agents.pragmatism_audit}β checks for over-engineering, god objects, premature abstractionshockey-analytics-smeβ if any stat counting, goal logic, or hockey domain is touched
Prompt format (adapt per specialist):
Review these planned tasks for epic [{milestone_id}] {epic_title}.
The epic goal: {goal}.
Tasks:
{numbered task list with Steps and Acceptance Criteria}
Flag: (1) missing files/functions we should read, (2) wrong approach for this domain,
(3) acceptance criteria that won't catch real failures, (4) scope that will cause issues.
Keep response under 300 words β actionable gaps only.
Incorporate any flagged gaps by updating task descriptions before creating them.
Document which specialists reviewed in the epic description under ## Specialist Review.
5f: Update Epic with Task References
After all tasks are created, update each epic's description to include the task {prefix}-{n} identifiers in the ### Tasks section:
save_issue({
id: "{epic_identifier}",
description: "...updated with task list..."
})
5g: Verify All Issues Linked Correctly (CRITICAL)
After creating all epics and tasks, run a verification sweep:
# Pull all issues just created
issues = list_issues(team: "{task_team_id}", query: "[{milestone_id}]")
# Verify each one
for issue in issues:
assert issue.project.id == {project_id}, f"{issue.identifier} has WRONG project: {issue.project.name}"
assert issue.milestone.id == {milestone_id_linear}, f"{issue.identifier} has WRONG milestone"
assert "milestone:{milestone_id}" in issue.labels, f"{issue.identifier} missing milestone label"
Print verification report:
LINEAR VERIFICATION
ββββββββββββββββββ
Issue Project Milestone Labels Status
ββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββββββ ββββββ
{prefix}-{n} β {project_name} β {milestone_name} β epic, milestone:... OK
{prefix}-{n} β {project_name} β {milestone_name} β task, milestone:... OK
{prefix}-{n} β WRONG: {actual} β {milestone_name} β missing milestone: FIX!
{pass_count}/{total_count} linked correctly
If ANY issues are wrong: fix them immediately with save_issue before proceeding.
If ALL pass: continue to 5h.
5h: Set Task Dependencies
Use Linear's blockedBy field:
save_issue({
id: "{task_identifier}",
blockedBy: ["{prev_task_identifier}"]
})
The REVIEW task is ALWAYS blocked by ALL other tasks in its epic.
Step 5b: Write Task Plans to Supabase
After all task-manager issues are created, record each task via save_task_plan (supabase β {storage_schema_qa}.task_plans) to enable PR review cross-checks:
For each task created in Step 5a (storage backend op; no-op under storage_backend: none):
save_task_plan({task_ID}, { milestone: "{milestone_id}", epic: "{epic_ID}",
goal: "{task goal}", steps, acceptance, planned_agent, planned_model, planned_machine,
status: "planned" })
This feeds the PR review loop β reviewers get_task_plan({task_ID}) to cross-check the implementation against the approved plan.
Step 6: Save Plan
-
The Linear project IS the plan. No separate plan files. No IMPLEMENTATION_PLAN.md updates.
-
Write decisions via
record_decision(storage) for any decisions made during planning -
Report β Generate a full milestone summary with execution diagram:
Milestone {milestone_id}: {goal} Epics: {count} | Tasks: {count} (all with parent epic relationships) EPIC ASSIGNMENTS | # | {prefix}-{id} | Title | Deps | Run Order | Primary | Backup | |---|----------|----------------------|------------|-----------|----------------------|----------------------| | 1 | {prefix}-{n} | VERIFY | none | 1 | workhorse/claude/son | auditor/gemini/flash | | 2 | {prefix}-{n} | {title} | VERIFY | 2 (parallel w/ E3) | workhorse/codex/gpt4 | auditor/gemini/flash | | 3 | {prefix}-{n} | {title} | VERIFY | 2 (parallel w/ E2) | auditor/gemini/flash | workhorse/codex/gpt4 | | 4 | {prefix}-{n} | {title} | E2 | 3 | mothership/cursor/son| workhorse/codex/gpt4 | | 5 | {prefix}-{n} | {title} | E2, E3 | 4 | workhorse/claude/son | auditor/gemini/flash | | 6 | {prefix}-{n} | HARDEN | ALL (1-5) | 5 | workhorse/claude/opus| mothership/gemini/pro| | 7 | {prefix}-{n} | Z: E2E | HARDEN | 6 | mothership/claude/op | workhorse/gemini/pro | Format: machine/agent/model (abbreviated to fit) Execution Plan VERIFY ({prefix}-{n}) / | \ E02 ({n}) E03 ({n}) E05 ({n}) β parallel (all depend only on VERIFY) | | E04 ({n}) | β depends on E02 \ / HARDEN ({n}) β depends on ALL above | Z ({n}) β depends on HARDEN Run Order: Step 1: VERIFY Step 2: E02, E03, E05 (parallel β different machines) Step 3: E04 (blocked by E02) Step 4: HARDEN (blocked by ALL) Step 5: Z (blocked by HARDEN) {count} Tasks: {prefix}-{first} through {prefix}-{last} Test Scaffold - tests/milestones/{milestone_id}/test_{milestone_id}_acceptance.py β {count} tests, all RED/skipped - {count} test classes: {class names with test counts} DoR compliance: all tasks passed self-check. Ready for factory: run `/epic start {prefix}-{first_epic}` to begin.Diagram rules:
- VERIFY is always the root node (no dependencies)
- Epics with no cross-dependencies are shown on the same level (β parallel)
- Epics with dependencies shown vertically with
|connectors - HARDEN always depends on ALL feature epics
- Z (if present) always depends on HARDEN
- Include {prefix}-{id} numbers so the diagram is actionable
- Use ASCII art β no unicode, no mermaid, just pipes and slashes
- Run Order section lists steps sequentially, noting which can run in parallel
- Parallel epics should be assigned to DIFFERENT machines when possible
Backup assignment rules (CRITICAL):
- Every epic gets Primary + Backup 1 + Backup 2 β three different agents
- Backup agent MUST be a different agent than primary AND each other (never claudeβclaudeβgemini)
- Backup machine SHOULD be a different machine than primary
- Backup model MUST be from the backup agent's ecosystem
- Ollama backups are ALWAYS on mothership (GPU required)
- If primary fails/is unavailable, factory falls back to Backup 1 automatically, then Backup 2
- Examples:
- Primary: claude/opus β B1: gemini/2.5-pro β B2: codex/gpt-4.1
- Primary: codex/gpt-4.1 β B1: ollama/qwen3:32b β B2: gemini/2.5-flash
- Primary: cursor/sonnet β B1: codex/gpt-4.1 β B2: ollama/qwen3:32b
- Primary: ollama/qwen3:32b β B1: codex/gpt-4.1 β B2: gemini/2.5-flash
- NEVER: claude/opus β claude/sonnet β gemini (same agent in primary + backup)
Milestone-level backup β shown at the top of the summary:
Milestone Backup (if Claude unavailable): workhorse/codex/gpt-4.1 Ollama fallback (always available): mothership/ollama/qwen3:32b
Step 7: Find or Create Master Tracker Issue
First: search for an existing tracker before creating.
results = search_issues(
query: "[{milestone_id}] π MASTER TRACKER",
filter: { label: "milestone-tracker" }
)
If tracker found (reopened milestone case):
- Read the existing tracker body to understand current state
- Update its body with current epic inventory (all epics, π² for not-started, β /β³ for any already done)
- Add a comment:
Milestone reopened {date}. Plan updated via /plan-milestone. Epics re-evaluated. - Note the existing
{task_prefix}-{tracker_number}β skip to "Set milestone description" below
If tracker NOT found (new milestone):
save_issue(
title: "[{milestone_id}] π MASTER TRACKER β {goal}",
team: "{task_team_id}",
project: "{project}",
milestone: "{milestone_linear_id}",
priority: 1,
labels: ["milestone-tracker", "milestone:{milestone_id}"],
description: {full tracker body β use Step 3b template from /milestone start skill}
)
Populate at plan time:
- Epic inventory table: all epics, all π², with blast radius + confidence π’ (freshly planned)
- Mermaid dependency graph: from the dependency ordering in Step 3b/3c
- Decisions table: decisions made during this planning session
- Session estimate: total sessions (never weeks β this is an AI factory)
- Acceptance criteria: from Step 2
- Health emoji: π’ (just planned)
Link tracker to every epic created:
for each epic {ID}:
save_issue(id: "{ID}", relatedTo: ["{tracker_id}"])
Set milestone Linear description (always β whether tracker was found or created):
**Goal:** {one sentence}
**Why:** {one sentence}
**Estimate:** {N} sessions
**Master tracker:** {task_prefix}-{tracker_number}
> Always reference this milestone as: **{milestone_id} / {task_prefix}-{tracker_number}**
Print: π Master tracker: {task_prefix}-{tracker_number} β always reference as {milestone_id} / {task_prefix}-{tracker_number}
Key Rules
- Acceptance criteria BEFORE epics. Epics BEFORE tasks. Tests BEFORE code.
- Every task must have testable completion criteria in the issue body
- Every task must pass the DoR self-check (5d.3) before creation
- Hardening epic is MANDATORY and ALWAYS last
- Dependencies must be explicit β no implicit ordering
- Epic size: 5-15 tasks. Bigger β split. Smaller β merge with related epic.
- Task size: S (< 30min), M (30min-2h), L (2h-4h). XL β split into subtasks.
- Labels drive factory routing:
machine:*,size:*,priority:*,type:* - Never create tasks for work that's already done (check closed issues first)
- Cross-reference spec sections β every task should trace back to a spec requirement
- Required Reading files must exist β verify with
test -fbefore including - Steps must be concrete β file paths, function names, specific changes (not "implement X")
Judgment weave (see /judgment)
Before finalizing the epic/task breakdown:
/premortemthe milestone plan β 5β8 past-tense causes of death; HΓH risks get redesigned or gated.- Every epic's acceptance criteria include at least one
GATE:line (metric | threshold | measured how | on-fail). "Works correctly" is not acceptance. - Any task that touches a one-way door (schema, stored formats, public contracts) is flagged with
/dooroutput in its ticket body, so the executor inherits the lock-in analysis instead of rediscovering it. - Plan-level verdicts β
/verdict log.