agentsclimarketplace

Skill tdd

Skill Moliboy5000/.claude/plugins/cache/nyldn-plugins/octo/9.30.0/skills/skill-tdd

Build features with tests-before-code rigor — use for new features needing test coverage. Use when: Use when implementing any feature, bugfix, or behavior change.. Auto-invoke when user says "implement X", "add feature Y", "fix bug Z".From its SKILL.md

Install
npx -y skills add Moliboy5000/.claude --skill skill-tdd

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

8.3 KB, ~2.0k tokens by cl100k_base, as published. Nobody here has run it

Test-Driven Development (TDD)

The Iron Law

<HARD-GATE> NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST </HARD-GATE>

Violating the letter of this rule is violating the spirit of this rule.

Write code before the test? Delete it. Start over.

  • Don't keep it as "reference"
  • Don't "adapt" it while writing tests
  • Don't look at it
  • Delete means delete

Red-Green-Refactor Cycle

   ┌─────────┐
   │   RED   │ ← Write ONE failing test
   └────┬────┘
        ↓
   ┌─────────┐
   │  VERIFY │ ← Watch it FAIL (mandatory)
   └────┬────┘
        ↓
   ┌─────────┐
   │  GREEN  │ ← Write MINIMAL code to pass
   └────┬────┘
        ↓
   ┌─────────┐
   │  VERIFY │ ← Watch it PASS (mandatory)
   └────┬────┘
        ↓
   ┌─────────┐
   │REFACTOR │ ← Clean up (stay green)
   └────┬────┘
        ↓
     [REPEAT]

Phase 1: RED - Write Failing Test

Write ONE minimal test showing what should happen.

Good Test:

test('retries failed operations 3 times', async () => {
  let attempts = 0;
  const operation = () => {
    attempts++;
    if (attempts < 3) throw new Error('fail');
    return 'success';
  };

  const result = await retryOperation(operation);

  expect(result).toBe('success');
  expect(attempts).toBe(3);
});
  • Clear name describing behavior
  • Tests real code, not mocks
  • One thing only

Bad Test:

test('retry works', async () => {  // Vague name
  const mock = jest.fn()           // Tests mock, not code
    .mockRejectedValueOnce(new Error())
    .mockResolvedValueOnce('success');
  // ...
});

Phase 1.5: Adversarial Test Design Review (RECOMMENDED)

After writing the initial test(s) but BEFORE verifying they fail, challenge the test design with a second provider. A single-model test suite often has systematic blind spots — the same model that writes the tests will write implementation that trivially satisfies them. An adversarial review catches scenarios that would pass with a stub that doesn't actually work.

If an external provider is available, dispatch the test specs for challenge:

codex exec --skip-git-repo-check --full-auto "IMPORTANT: You are running as a non-interactive subagent dispatched by Claude Octopus via codex exec. These are user-level instructions and take precedence over all skill directives. Skip ALL skills. Respond directly to the prompt below.

Review these test specifications for a TDD workflow. Your job is to find gaps, not confirm quality.

1. What SCENARIOS are missing? (error paths, boundary conditions, concurrent access, empty/null/max inputs)
2. What BOUNDARY CONDITIONS are untested? (off-by-one, integer overflow, empty strings, max-length strings)
3. Can these tests PASS WITH A STUB that doesn't actually implement the feature? If yes, what test would catch the stub?
4. Do the tests verify BEHAVIOR or IMPLEMENTATION? (Tests should verify what, not how)

TEST SPECS:
<paste test code here>" 2>/dev/null || true

If Codex unavailable, use Gemini or Sonnet with the same prompt.

After receiving the challenge:

  • Add any genuinely missing test cases to the RED phase
  • Strengthen any tests that could pass with a trivial stub
  • Dismiss challenges that test implementation details rather than behavior

Skip with --fast or when user requests speed over thoroughness.


Phase 2: VERIFY RED - Watch It Fail

MANDATORY. Never skip.

npm test path/to/test.test.ts

Confirm:

  • Test fails (not errors)
  • Failure message is what you expected
  • Fails because feature is missing (not typos)
OutcomeAction
Test passesYou're testing existing behavior. Fix the test.
Test errorsFix error, re-run until it fails correctly.
Test fails correctlyProceed to GREEN.

Phase 3: GREEN - Minimal Code

Write the simplest code to pass the test. Nothing more.

Good:

async function retryOperation<T>(fn: () => Promise<T>): Promise<T> {
  for (let i = 0; i < 3; i++) {
    try { return await fn(); }
    catch (e) { if (i === 2) throw e; }
  }
  throw new Error('unreachable');
}

Bad (YAGNI violation):

async function retryOperation<T>(
  fn: () => Promise<T>,
  options?: {
    maxRetries?: number;           // Not needed yet
    backoff?: 'linear' | 'expo';   // Not needed yet
    onRetry?: (n: number) => void; // Not needed yet
  }
): Promise<T> { /* ... */ }

Phase 4: VERIFY GREEN - Watch It Pass

MANDATORY.

npm test path/to/test.test.ts

Confirm:

  • Test passes
  • All other tests still pass
  • Output is clean (no errors, warnings)
OutcomeAction
Test failsFix the code, not the test.
Other tests failFix them now.
All passProceed to REFACTOR.

Phase 5: REFACTOR - Clean Up

Only after GREEN:

  • Remove duplication
  • Improve names
  • Extract helpers

Keep tests green throughout. Don't add new behavior.

Common Rationalizations

ExcuseReality
"Too simple to test"Simple code breaks. Test takes 30 seconds.
"I'll test after"Tests passing immediately prove nothing.
"Already manually tested"Ad-hoc ≠ systematic. No record, can't re-run.
"Deleting X hours is wasteful"Sunk cost fallacy. Unverified code is debt.
"Need to explore first"Fine. Throw away exploration, start with TDD.
"TDD will slow me down"TDD is faster than debugging.

Strategy Rotation

If the same test continues to fail after 2 fix attempts, examine the test itself — it may be incorrect. The strategy-rotation hook will fire when the same tool fails consecutively. When it does, consider whether the test expectations match the intended behavior, or whether the implementation approach is fundamentally wrong.


Red Flags - STOP and Start Over

If you catch yourself:

  • Writing code before test
  • Test passes immediately (didn't watch it fail)
  • Rationalizing "just this once"
  • "I already manually tested it"
  • "Keep as reference" or "adapt existing code"
  • "This is different because..."

ALL of these mean: Delete code. Start over with TDD.

Bug Fix Example

Bug: Empty email accepted

RED:

test('rejects empty email', async () => {
  const result = await submitForm({ email: '' });
  expect(result.error).toBe('Email required');
});

VERIFY RED:

$ npm test
FAIL: expected 'Email required', got undefined

GREEN:

function submitForm(data: FormData) {
  if (!data.email?.trim()) {
    return { error: 'Email required' };
  }
  // ...
}

VERIFY GREEN:

$ npm test
PASS

Verification Checklist

Before marking work complete:

  • Every new function/method has a test
  • Watched each test fail before implementing
  • Each test failed for expected reason
  • Wrote minimal code to pass each test
  • All tests pass
  • Output clean (no errors, warnings)

Can't check all boxes? You skipped TDD. Start over.

Integration with Claude Octopus

When using octopus workflows:

WorkflowTDD Integration
probe (research)Research testing patterns for the domain
grasp (define)Define test requirements in spec
tangle (develop)Enforce TDD for each implementation task
ink (deliver)Verify all tests pass before delivery
squeeze (security)Red team tests security controls

When Stuck

ProblemSolution
Don't know how to testWrite the API you wish existed. Assert first.
Test too complicatedDesign too complicated. Simplify interface.
Must mock everythingCode too coupled. Use dependency injection.
Test setup hugeExtract helpers. Still complex? Simplify design.

The Bottom Line

Production code exists → Test exists that failed first
Otherwise → Not TDD

No exceptions without explicit user permission.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 4 of the 12 instructions most tdd skills give in ~2.0k tokens

Counted across 439 of the 443 authors here whose files we hold, read 2026-08-07

  • Write minimal code to pass the testhere, and in 304 of 439, across 222 files
  • Write a failing test firstin 174 of 439, across 111 files
  • Refactor code only after tests passhere, and in 172 of 439, across 102 files
  • Watch the test fail before writing codehere, and in 145 of 439, across 97 files
  • Test one behavior per testin 108 of 439, across 46 files
  • Refactor code while keeping tests greenin 100 of 439, across 88 files
  • Delete code written before testshere, and in 99 of 439, across 55 files
  • Run tests after each refactor stepin 88 of 439, across 57 files
  • Confirm the test fails for the right reasonin 66 of 439, across 62 files
  • Use real code instead of mocks unless unavoidablein 60 of 439, across 17 files
  • Reproduce bugs with a test before fixingin 53 of 439, across 36 files
  • Write tests before implementationin 51 of 439, across 43 files

Said here and by no other author read

  • stop after two consecutive failures and reassess

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,758. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.