agentsclimarketplace

Flaky test investigation

Skill yeaight7/agent-powerups/skills/flaky-test-investigation

Use when tests pass and fail intermittently without code changes, or a test passes alone but fails in the full suite.From its SKILL.md

Install
npx -y skills add yeaight7/agent-powerups --skill flaky-test-investigation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

2.5 KB, 580 tokens by cl100k_base, as published. Nobody here has run it

Purpose

Flaky tests erode trust in CI. Do not just re-run them and hope for the best — isolate the flake vector, fix it, and prove the fix with a stress loop.

When to Use

  • A test fails intermittently in CI but passes locally (or vice versa)
  • A test passes alone but fails in the full suite
  • A re-run "fixed" a failure and nobody knows why

Inputs

  • The flaky test's name/path and the runner command for it
  • Recent failing runs, if available, to estimate the failure rate

Workflow

  1. Isolate the test. Run the specific failing test by itself. If it passes alone, the flake is likely an order dependency or state leakage from a previous test — run the suite up to and including it to confirm.

  2. Stress test. Run the test in a tight loop to establish the failure rate before changing anything:

    for i in {1..100}; do npm test -- -t "My Test" || echo "FAIL on run $i"; done
    

    (Adapt the inner command to the project's runner; some runners have repeat flags built in.)

  3. Check the common vectors:

    • Time — does the test rely on Date.now() or setTimeout? Mock the clock.
    • Async/Promises — asserting before a background task finishes? Ensure proper await or waitFor usage.
    • Shared state — reusing database records, global singletons, or mutated variables between runs? Ensure clean teardowns in afterEach.
    • Randomness — random IDs or sort orders? Force deterministic seeds or sort orders.
  4. Prove the fix. Do not just guess. The fix must be verified by running the stress test loop again and achieving a 100% pass rate.

Output

  • The identified flake vector (order/state, time, async, randomness)
  • The fix, plus stress-loop evidence (pre-fix failure rate vs post-fix 100% pass)

Verification

  • Test run in isolation to separate order-dependency from intrinsic flake
  • Stress loop run before the fix to establish a baseline failure rate
  • Flake vector named explicitly
  • Stress loop re-run after the fix with a 100% pass rate

Failure Modes

  • Re-run and hope — a green re-run proves nothing; the flake is still there.
  • Fixing without a baseline — without a pre-fix failure rate, a "fix" cannot be distinguished from luck.
  • Quarantining forever — skipping the test removes the signal but keeps the bug.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,852. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.