Testing
Use when writing tests, deciding on test coverage, choosing between unit/integration/E2E tests, applying TDD, or reviewing test quality — covers rigor tiers, the Red-Green-Refactor cycle, coverage requirements, and Playwright for UI testing.From its SKILL.md
npx -y skills add andr-ca/agentharness --skill testingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
3.6 KB, 808 tokens by cl100k_base, as published. Nobody here has run it
Testing
This skill is self-contained for day-to-day use. Deeper reference (needs
the full harness checkout): patterns/testing/TDD.md (full TDD
workflow), patterns/testing/COVERAGE_REQUIREMENTS.md (80% rule and
pragma), patterns/testing/PLAYWRIGHT_UI_TESTING.md (E2E screenshots),
patterns/testing/COMPLETION_CHECKLIST.md.
Rigor tiers — apply the right standard, not the maximum
| Tier | Coverage | TDD | Notes |
|---|---|---|---|
| Production | ≥ 80% line+branch | Required | Libraries, agents, shipped services |
| Internal | Must have tests | Recommended | No numeric floor; cover what's expensive to get wrong |
| Prototype | None required | Optional | Skip freely; document the decision |
Check .agentharness-profile or .github/CODING_GUIDELINES.md for the
project's tier before assuming 80% applies.
Red-Green-Refactor
- Red — write a failing test for the one behavior you're adding.
- Green — write the minimum code to make it pass (not the final code — the correct code comes in the next step).
- Refactor — clean up duplication and naming while tests stay green.
Never skip Red — a test that never fails doesn't prove anything.
Test one behavior per test
# WRONG: five assertions for one behavior — one failure hides the rest
def test_user_creation():
user = User.create(email="[email protected]", password="s3cr3t")
assert user.id is not None
assert user.email == "[email protected]"
assert user.password_hash is not None
assert user.created_at is not None
assert user.is_active is True
# RIGHT: one assertion covers "all fields are set" as a single behavior
def test_user_creation_populates_all_fields():
user = User.create(email="[email protected]", password="s3cr3t")
assert user == {
"id": AnyString(),
"email": "[email protected]",
"password_hash": AnyString(),
"created_at": AnyDatetime(),
"is_active": True,
}
Split into multiple tests only when the behaviors can independently fail for unrelated reasons.
Test pyramid shape (guide, not hard rule)
~70% unit · ~25% integration · ~5% E2E. Bias toward cheap-to-run unit tests; don't duplicate unit coverage with integration tests that only exist for the coverage number.
Coverage floor: 80% line + branch at Production tier
# pragma: no cover is for truly unreachable code — not for skipping inconvenient tests
if sys.version_info >= (3, 12):
new_behavior()
else: # pragma: no cover — never reached on supported versions
old_behavior()
Measure coverage before marking work complete. Don't approve a PR that claims ≥ 80% without a CI report to confirm it.
UI testing (Playwright)
For UI tests: screenshot-review every visual path at least once (not
just "it didn't crash"). State in the PR description which flows you
verified visually — CI can check that the test ran, but not that the UI
looked correct. See patterns/testing/PLAYWRIGHT_UI_TESTING.md for full
conventions.
Assertion style
Prefer specific assertions that produce readable failure messages:
// POOR: tells you nothing on failure except "false"
expect(result).toBe(true);
// GOOD: tells you what was in result
expect(result.status).toBe('active');
expect(result.items).toHaveLength(3);
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.