Running tdd cycles
Drive strict red-green-refactor TDD discipline on any code change, any language.From its SKILL.md
npx -y skills add swell-agents/coding-skills --skill running-tdd-cyclesAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.1 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it
Loop
For each new piece of behaviour:
1. Extract <requirement> — the smallest piece of logic that adds value.
2. RED → write ONE failing test that pins down <requirement>.
3. GREEN → write the minimal code that makes it pass.
4. REFACTOR → improve structure with the test as a safety net.
5. COMMIT → one logical change per commit (defer to committing-changes).
6. Repeat 2–5 until the task is done.
7. REVIEW → defer to reviewing-changes for the final pass.
Always one requirement per cycle. If the cycle feels big, the requirement was too big — split it.
RED — write a failing test
- One test, one requirement. Don't write a test suite up front; one test, fail, pass, refactor, next.
- Test names are claims, not labels. Verify the fixture shape, input source, and assertion target match the name before the test goes green. A test named
test_decodesV4Quotethat loads a V5 fixture is silently broken: it passes but pins the wrong invariant, and the regression lands in the wrong place. Common causes: renamed a test and forgot to update the fixture; copy-pasted a fuzzy assertion from a sibling test without re-targeting it. - Arrange-Act-Assert. Three blocks, one assertion focus. Use
should_X_when_Ynaming where the framework allows it. - Fail for the right reason. Run the test before writing implementation. The failure message must point at missing behaviour, not at a typo, missing import, or fixture mistake. If it doesn't, the test is wrong.
- No premature edge cases. Cover the happy path first. Edge cases (
null, empty collections, boundary values, concurrent access, errors) get their own RED-GREEN-REFACTOR cycles. - Property-based when applicable. Prefer
hypothesis(Python),fast-check(JS/TS), orquickcheck(Haskell/Rust) for invariants. Example-based tests are for specific edge cases or documentation. Seereference/property-based-testing.md. - Reuse strategies. Before writing a generator, check
tests/strategies.py(or the project equivalent) for shared fixtures/strategies. Add new ones back to the shared module.
Framework cues
- Python (pytest) — fixtures with explicit scopes,
pytest.mark.parametrizefor multiple cases,hypothesisfor properties. - JS/TS (Vitest/Jest) —
vi.fn()/jest.fn()mocks;@testing-library/*for components;fast-checkfor properties. - Go — table-driven tests with subtests via
t.Run(name, func(t *testing.T) {...});t.Parallel()where safe. - Ruby (RSpec) —
letfor lazy fixtures,let!for eager; nestedcontextfor branching scenarios. - Solidity (Foundry) —
forge test -vvv; parametric fuzz tests viafunction testFoo(uint256 x);bound(x, min, max)notif-guards. See solidity-conventions.
GREEN — minimal code to pass
- Smallest possible change. Hard-coded returns are fine for the first cycle. Triangulate (generalise) only when a second test forces it.
- No bonus features. Don't add error handling, logging, optimisation, or generality unless a test demands it. The next cycle will.
- Don't modify the test. If the test is wrong, go back to RED. If the test is right and the implementation is wrong, fix the implementation.
- Run the full test suite after each change. Confirm green; confirm no regression.
REFACTOR — improve structure
- Tests stay green throughout. Run the suite after every micro-step.
- Code smells to act on: duplication (extract method/class), long methods (decompose), large classes (split SRP), long parameter lists (parameter object), feature envy (move method), primitive obsession (value object), switch statements driven by type (polymorphism), dead code (delete).
- SOLID weights. Single Responsibility first; the others when they pull their weight. See engineering-philosophy.
- Tests refactor too. Extract common fixtures, rename for clarity, eliminate test duplication. Coverage stays equal or improves.
- Performance refactors are measured. Profile before, profile after; commit the measurement alongside the change.
- Refactoring is not optional. If you skip it, technical debt compounds and the next RED gets harder to write.
Anti-patterns
- Writing implementation before the test.
- Writing a test that already passes.
- Writing many tests at once and implementing them in a batch.
- Modifying a test to make it pass.
- Skipping refactor because "the test passed."
- Adding tests during the GREEN phase.
- Test-after rationalised as TDD.
If the discipline breaks: stop, identify the violated phase, revert to the last green state, resume from the right phase, note what went wrong.
Validation checkpoints
End of RED: test exists, test fails, failure message names the missing behaviour, no false positive.
End of GREEN: all tests pass, the change is the minimum that could possibly work, the test wasn't modified, coverage didn't drop.
End of REFACTOR: all tests still pass, complexity (cyclomatic, length) at least no worse, duplication addressed, names readable, performance measured if performance was the goal.
Scratch testing
Quick exploration belongs in a gitignored scratch file, not in production code or in the test suite:
- Python →
test.py(see python-conventions andreference/scratch-testing.md). - Go →
test.gowith//go:build scratchtag (see go-conventions). - Solidity →
script/Scratch.s.solorchiselREPL (see solidity-conventions).
Never use inline heredocs (uv run python << EOF, go run -, forge script --via-stdin) — the file pattern keeps history and stays inspectable.
Cross-references
committing-changes— commit after every successful GREEN or REFACTOR phase.reviewing-changes— final pass after the loop ends.python-conventions/go-conventions/solidity-conventions— language-specific test runner setup.engineering-philosophy— Small Steps, Investigate-Don't-Mask, KISS, YAGNI all apply directly.
Reference
- reference/cycle-deep-dive.md — full 12-phase pipeline with coverage thresholds and metrics.
- reference/tdd-orchestrator-agent.md — original Claude-Code orchestrator agent verbatim (model: opus).
- reference/test-automator-agent.md — original Claude-Code test-automator agent verbatim (model: sonnet).
- reference/property-based-testing.md — Hypothesis, fast-check, QuickCheck patterns.
- reference/scratch-testing.md — language-specific scratch-file patterns.
- reference/tdd-workflow-rule.md — original
rules/tdd-workflow.mdverbatim.
What ships with it: 6 files
34.1 KB alongside SKILL.md
reference/
- cycle-deep-dive.md3.7 KB
- property-based-testing.md4.1 KB
- scratch-testing.md2.3 KB
- tdd-orchestrator-agent.md11.7 KB
- tdd-workflow-rule.md852 B
- test-automator-agent.md11.5 KB
Gives 3 of the 12 instructions most tdd skills give in ~1.6k tokens
Counted across 594 of the 695 authors here whose files we hold, read 2026-09-06
- Write minimal code to pass the testhere, and in 356 of 594, across 332 files
- Write a failing test before writing production codein 240 of 594, across 221 files
- Refactor code only after tests passin 184 of 594, across 169 files
- Refactor code while keeping tests greenhere, and in 140 of 594, across 133 files
- Verify the test fails for the expected reasonin 133 of 594, across 121 files
- Run tests after each refactor stepin 108 of 594, across 102 files
- Use real code instead of mocks whenever possiblein 97 of 594, across 85 files
- Reproduce bugs with a failing test before fixingin 90 of 594, across 81 files
- Write tests before implementing codein 88 of 594, across 70 files
- Verify the test passes after writing codein 71 of 594, across 61 files
- Write one test for one behaviorin 71 of 594, across 63 files
- Run the full test suitehere, and in 63 of 594, across 56 files
Said here and by no other author read
- Extract the smallest requirement adding value
- Commit one logical change per cycle
- Split requirements if the cycle feels too big
- Verify test failure points to missing behaviour
- Profile performance before and after refactoring
- Use scratch files for quick exploration
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.