Tdd workflow
Use this skill when writing new features, fixing bugs, or refactoring code. Enforces test-driven development with around 90% line coverage of real logic including unit, integration, and E2E tests.From its SKILL.md
npx -y skills add Lukk17/agent-standards --skill tdd-workflowAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
16.8 KB, ~4.1k tokens by cl100k_base, as published. Nobody here has run it
Test-Driven Development Workflow
This skill ensures all code development follows TDD principles with comprehensive test coverage.
When to Activate
- Writing new features or functionality
- Fixing bugs or issues
- Refactoring existing code
- Adding API endpoints
- Creating new components
Core Principles
1. Tests BEFORE Code
ALWAYS write tests first, then implement code to make tests pass.
2. Coverage Requirements
- Around 90% line coverage of real logic (unit + integration + E2E); branch coverage may stay at 70 to 80%
- 100% for critical logic where it genuinely adds value
- No exclusion patterns to dodge meaningful tests
- All edge cases covered
- Error scenarios tested
- Boundary conditions verified
3. Test Types
Unit Tests
- Individual functions and utilities
- Component logic
- Pure functions
- Helpers and utilities
Integration Tests
- API endpoints
- Database operations
- Service interactions
- External API calls
E2E Tests (Playwright)
- Critical user flows
- Complete workflows
- Browser automation
- UI interactions
TDD Workflow Steps
Step 1: Write User Journeys
As a [role], I want to [action], so that [benefit]
Example:
As a user, I want to search for markets semantically,
so that I can find relevant markets even without exact keywords.
Step 2: Generate Test Cases
For each user journey, create comprehensive test cases:
describe('Semantic Search', () => {
it('returns relevant markets for query', async () => {
// Test implementation
})
it('handles empty query gracefully', async () => {
// Test edge case
})
it('falls back to substring search when Redis unavailable', async () => {
// Test fallback behavior
})
it('sorts results by similarity score', async () => {
// Test sorting logic
})
})
Step 3: Run Tests (They Should Fail)
npm test
# Tests should fail - we haven't implemented yet
This step is mandatory and is the RED gate for all production changes.
Before modifying business logic or other production code, you must verify a valid RED state via one of these paths:
- Runtime RED:
- The relevant test target compiles successfully
- The new or changed test is actually executed
- The result is RED
- Compile-time RED:
- The new test newly instantiates, references, or exercises the buggy code path
- The compile failure is itself the intended RED signal
- In either case, the failure is caused by the intended business-logic bug, undefined behavior, or missing implementation
- The failure is not caused only by unrelated syntax errors, broken test setup, missing dependencies, or unrelated regressions
A test that was only written but not compiled and executed does not count as RED.
Do not edit production code until this RED state is confirmed.
Step 4: Implement Code
Write minimal code to make tests pass:
// Implementation guided by tests
export async function searchMarkets(query: string) {
// Implementation here
}
Step 5: Run Tests Again
npm test
# Tests should now pass
Rerun the same relevant test target after the fix and confirm the previously failing test is now GREEN.
Only after a valid GREEN result may you proceed to refactor.
Step 6: Refactor
Improve code quality while keeping tests green:
- Remove duplication
- Improve naming
- Optimize performance
- Enhance readability
Step 7: Verify Coverage
npm run test:coverage
# Verify ~90% line coverage of real logic achieved
Testing Patterns
Unit Test Pattern (Jest/Vitest)
import { render, screen, fireEvent } from '@testing-library/react'
import { Button } from './Button'
describe('Button Component', () => {
it('renders with correct text', () => {
render(<Button>Click me</Button>)
expect(screen.getByText('Click me')).toBeInTheDocument()
})
it('calls onClick when clicked', () => {
const handleClick = jest.fn()
render(<Button onClick={handleClick}>Click</Button>)
fireEvent.click(screen.getByRole('button'))
expect(handleClick).toHaveBeenCalledTimes(1)
})
it('is disabled when disabled prop is true', () => {
render(<Button disabled>Click</Button>)
expect(screen.getByRole('button')).toBeDisabled()
})
})
API Integration Test Pattern
import { NextRequest } from 'next/server'
import { GET } from './route'
describe('GET /api/markets', () => {
it('returns markets successfully', async () => {
const request = new NextRequest('http://localhost/api/markets')
const response = await GET(request)
const data = await response.json()
expect(response.status).toBe(200)
expect(data.success).toBe(true)
expect(Array.isArray(data.data)).toBe(true)
})
it('validates query parameters', async () => {
const request = new NextRequest('http://localhost/api/markets?limit=invalid')
const response = await GET(request)
expect(response.status).toBe(400)
})
it('handles database errors gracefully', async () => {
// Mock database failure
const request = new NextRequest('http://localhost/api/markets')
// Test error handling
})
})
E2E Test Pattern (Playwright)
import { test, expect } from '@playwright/test'
test('user can search and filter markets', async ({ page }) => {
// Navigate to markets page
await page.goto('/')
await page.click('a[href="/markets"]')
// Verify page loaded
await expect(page.locator('h1')).toContainText('Markets')
// Search for markets
await page.fill('input[placeholder="Search markets"]', 'election')
// Wait for debounce and results
await page.waitForTimeout(600)
// Verify search results displayed
const results = page.locator('[data-testid="market-card"]')
await expect(results).toHaveCount(5, { timeout: 5000 })
// Verify results contain search term
const firstResult = results.first()
await expect(firstResult).toContainText('election', { ignoreCase: true })
// Filter by status
await page.click('button:has-text("Active")')
// Verify filtered results
await expect(results).toHaveCount(3)
})
test('user can create a new market', async ({ page }) => {
// Login first
await page.goto('/creator-dashboard')
// Fill market creation form
await page.fill('input[name="name"]', 'Test Market')
await page.fill('textarea[name="description"]', 'Test description')
await page.fill('input[name="endDate"]', '2025-12-31')
// Submit form
await page.click('button[type="submit"]')
// Verify success message
await expect(page.locator('text=Market created successfully')).toBeVisible()
// Verify redirect to market page
await expect(page).toHaveURL(/\/markets\/test-market/)
})
Test File Organization
src/
├── components/
│ ├── Button/
│ │ ├── Button.tsx
│ │ ├── Button.test.tsx # Unit tests
│ │ └── Button.stories.tsx # Storybook
│ └── MarketCard/
│ ├── MarketCard.tsx
│ └── MarketCard.test.tsx
├── app/
│ └── api/
│ └── markets/
│ ├── route.ts
│ └── route.test.ts # Integration tests
└── e2e/
├── markets.spec.ts # E2E tests
├── trading.spec.ts
└── auth.spec.ts
Mocking External Services
The mocks below are for fast unit tests only. Do not mock the database when you can exercise it: integration tests must use Testcontainers or a real engine, as mandated in the "Testcontainers Mandate" and "Do not mock what you do not own" sections later in this skill. The Supabase and Redis mocks here stand in for an external boundary in a unit test; they are not a substitute for exercising the real datastore in integration tests.
Supabase Mock (unit tests only)
jest.mock('@/lib/supabase', () => ({
supabase: {
from: jest.fn(() => ({
select: jest.fn(() => ({
eq: jest.fn(() => Promise.resolve({
data: [{ id: 1, name: 'Test Market' }],
error: null
}))
}))
}))
}
}))
Redis Mock
jest.mock('@/lib/redis', () => ({
searchMarketsByVector: jest.fn(() => Promise.resolve([
{ slug: 'test-market', similarity_score: 0.95 }
])),
checkRedisHealth: jest.fn(() => Promise.resolve({ connected: true }))
}))
OpenAI Mock
jest.mock('@/lib/openai', () => ({
generateEmbedding: jest.fn(() => Promise.resolve(
new Array(1536).fill(0.1) // Mock 1536-dim embedding
))
}))
Test Coverage Verification
Run Coverage Report
npm run test:coverage
Coverage Thresholds
{
"jest": {
"coverageThresholds": {
"global": {
"branches": 70,
"functions": 90,
"lines": 90,
"statements": 90
}
}
}
}
Common Testing Mistakes to Avoid
FAIL: WRONG: Testing Implementation Details
// Don't test internal state
expect(component.state.count).toBe(5)
PASS: CORRECT: Test User-Visible Behavior
// Test what users see
expect(screen.getByText('Count: 5')).toBeInTheDocument()
FAIL: WRONG: Brittle Selectors
// Breaks easily
await page.click('.css-class-xyz')
PASS: CORRECT: Semantic Selectors
// Resilient to changes
await page.click('button:has-text("Submit")')
await page.click('[data-testid="submit-button"]')
FAIL: WRONG: No Test Isolation
// Tests depend on each other
test('creates user', () => { /* ... */ })
test('updates same user', () => { /* depends on previous test */ })
PASS: CORRECT: Independent Tests
// Each test sets up its own data
test('creates user', () => {
const user = createTestUser()
// Test logic
})
test('updates user', () => {
const user = createTestUser()
// Update logic
})
Continuous Testing
Watch Mode During Development
npm test -- --watch
# Tests run automatically on file changes
Pre-Commit Hook
# Runs before every commit
npm test && npm run lint
CI/CD Integration
# GitHub Actions
- name: Run Tests
run: npm test -- --coverage
- name: Upload Coverage
uses: codecov/codecov-action@v3
Best Practices
- Write a failing test first - Always TDD
- One behavior per test - Focus on a single outcome; multiple assertions are fine when they all verify that one outcome
- Descriptive Test Names - Explain what's tested
- Given / When / Then - Split every test body into three labelled sections with
// Given,// When, and// Thencomments (Given sets up state, When runs the single action, Then asserts the observable outcome) - Mock only external dependencies - Isolate unit tests; do not mock the database when you can exercise it
- Test Edge Cases - Null, undefined, empty, large
- Test Error Paths - Not just happy paths
- Never weaken an assertion - Fix the code or the test setup; do not loosen a check to make a test pass
- Keep Tests Fast - Unit tests < 50ms each
- Clean Up After Tests - No side effects
- Review Coverage Reports - Identify gaps
Success Metrics
- Around 90% line coverage of real logic achieved (100% for critical logic where it adds value)
- All tests passing (green)
- No skipped or disabled tests
- Fast test execution (< 30s for unit tests)
- E2E tests cover critical user flows
- Tests catch bugs before production
Remember: Tests are not optional. They are the safety net that enables confident refactoring, rapid development, and production reliability.
Advanced Testing Standards
Test Naming Convention
Use the <methodName>_<scenario>_<expectedResult> pattern:
// Java (JUnit 5)
@Test
void calculateTotal_withEmptyCart_returnsZero() { }
@Test
void processPayment_whenCardDeclined_throwsPaymentException() { }
// TypeScript (Vitest)
test('calculateTotal_withEmptyCart_returnsZero', () => { })
test('processPayment_whenCardDeclined_throwsPaymentException', () => { })
Never use vague names like test1, works, or happyPath.
TestDataFactory Pattern
Never repeat object construction inline across tests. Extract to a shared factory:
// Java
public class OrderTestFactory {
public static Order validOrder() {
return Order.builder()
.id(UUID.randomUUID())
.customerId("cust-001")
.items(List.of(OrderItem.of("SKU-1", 2, BigDecimal.valueOf(9.99))))
.status(OrderStatus.PENDING)
.build();
}
public static Order cancelledOrder() {
return validOrder().toBuilder().status(OrderStatus.CANCELLED).build();
}
}
// TypeScript
export const OrderFactory = {
valid: (): Order => ({
id: crypto.randomUUID(),
customerId: 'cust-001',
items: [{ sku: 'SKU-1', qty: 2, price: 9.99 }],
status: 'pending',
}),
cancelled: (): Order => ({ ...OrderFactory.valid(), status: 'cancelled' }),
}
Test Pyramid
Maintain these proportions across the test suite:
| Layer | Target | Tools |
|---|---|---|
| Unit | ~70% | JUnit/Vitest, fast, no I/O |
| Integration | ~20% | Testcontainers, real DB/queue |
| Contract | ~5% | Pact, Spring Cloud Contract |
| E2E | ~5% | Playwright, Cypress, Selenium |
Contract Testing (Required for External HTTP APIs)
Every HTTP API consumed by an external service must have consumer-driven contract tests:
// Spring Cloud Contract (provider side)
Contract.make {
request {
method 'GET'
url '/api/users/123'
}
response {
status 200
body([id: '123', name: 'Alice'])
headers { contentType(applicationJson()) }
}
}
Use Pact for polyglot environments; Spring Cloud Contract for Java-to-Java service contracts.
Mutation Testing
Run mutation testing on all critical business logic:
- Java: PIT (
pitest): minimum 70% mutation score as CI gate - TypeScript/JavaScript: Stryker: minimum 70% mutation score as CI gate
<!-- Java pom.xml -->
<plugin>
<groupId>org.pitest</groupId>
<artifactId>pitest-maven</artifactId>
<configuration>
<mutationThreshold>70</mutationThreshold>
<coverageThreshold>80</coverageThreshold>
</configuration>
</plugin>
Performance Testing
Required before major releases and for any change to a hot path:
- k6 (HTTP load testing) or Gatling (JVM) for service endpoints
- pytest-benchmark for Python critical paths
- Alert on p99 regression > 20% versus the previous release baseline
// k6 example
export const options = {
thresholds: {
http_req_duration: ['p(99)<200'], // p99 must be under 200ms
http_req_failed: ['rate<0.01'], // Error rate < 1%
},
}
Flaky Test Policy
Flaky tests are classified as blocking defects:
| Rule | Value |
|---|---|
| Fix or quarantine SLA | 2 business days |
| Maximum quarantine period | 2 sprints |
| Action after quarantine expires | Delete the test (rewrite from scratch) |
| Prohibited patterns | Thread.sleep(), time.sleep(), setTimeout in test assertions |
| Allowed retry | @RetryingTest (JUnit) only for inherently non-deterministic integration tests |
Coverage Thresholds
Enforced as a CI gate, PRs that drop coverage below threshold are blocked:
- Line coverage: minimum 90% of real logic (100% for critical logic where it genuinely adds value)
- Branch coverage: minimum 70% (70 to 80% is acceptable)
<!-- JaCoCo Maven config -->
<rule>
<element>BUNDLE</element>
<limits>
<limit>
<counter>LINE</counter>
<value>COVEREDRATIO</value>
<minimum>0.90</minimum>
</limit>
<limit>
<counter>BRANCH</counter>
<value>COVEREDRATIO</value>
<minimum>0.70</minimum>
</limit>
</limits>
</rule>
Key Rules
Do not mock what you do not own. Only mock types you define. For third-party libraries (HTTP clients, ORMs, cloud SDKs), use:
- Real instances via Testcontainers
- Official test doubles provided by the library
- WireMock / MSW for HTTP boundaries
Run the full test suite, never run a single test in isolation to verify a fix. A fix that makes one test pass but breaks another is not a fix.
Testcontainers Mandate
All integration tests that touch external systems (databases, message queues, caches, cloud services) must use Testcontainers:
@Testcontainers
class OrderRepositoryTest {
@Container
static PostgreSQLContainer<?> postgres = new PostgreSQLContainer<>("postgres:16-alpine");
@DynamicPropertySource
static void props(DynamicPropertyRegistry registry) {
registry.add("spring.datasource.url", postgres::getJdbcUrl);
}
}
Never use a shared staging database for automated tests, tests must be hermetic and reproducible.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.