Fix flaky test
Use when a ticket reports an intermittently failing test — passes sometimes, fails others, or fails in CI but not locally. Invoke for "this test is flaky", "fix the intermittent failure in X", or "the suite is non-deterministic". Fix the root cause, never paper over it with a retry.From its SKILL.md
npx -y skills add tmj-90/gaffer --skill fix-flaky-testAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
2.0 KB, 421 tokens by cl100k_base, as published. Nobody here has run it
Fix a flaky test
A flaky test fails non-deterministically. The job is to find the source of non-determinism and remove it — not to retry until it passes.
Steps
- Reproduce the flakiness. Run the named test repeatedly (a loop, or the runner's repeat flag) and, where relevant, in randomised order. Capture a failing run; you can't fix what you can't observe.
- Locate the non-determinism. It's almost always one of: timing (sleeps, races, unawaited async), test order / shared mutable state (leaked globals, DB rows, singletons), unseeded randomness, real clock/timezone, or network/external calls. Read the failure to narrow which.
- Fix the root cause. Await async properly and wait on conditions not timeouts; isolate state with proper setup/teardown; seed randomness and fake the clock; stub the external boundary. Do not add a retry, increase a sleep, or mark the test skipped — those hide the bug.
- Prove stability. Re-run the test many times (and in random order) and confirm it
passes every time. Then run the full suite (
run-tests) to confirm no regression. - Evidence: the repeated-run command and a clean streak, plus the fix summary. Then
use the
record-evidenceskill to recordtest_outputagainst the AC and submit for review.
Rules
- Fix the cause (timing/order/shared state/randomness), never add a retry or longer sleep.
- Never skip or delete the test to make the suite green.
- Demonstrate stability with many repeated passes, not a single run.
- If the flake reveals a real product bug, note it; fix only what the ticket scopes.
- Run on a branch (the
create-branchskill), never a protected branch.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most test skills give in 421 tokens
Counted across 964 of the 1,571 authors here whose files we hold, read 2026-08-07
- Close the browser when donein 55 of 964, across 12 files
- Wait for network idle statein 51 of 964, across 6 files
- Launch Chromium in headless modein 49 of 964, across 6 files
- Use descriptive selectors for elementsin 49 of 964, across 6 files
- Run provided scripts with help flag firstin 49 of 964, across 6 files
- Add appropriate explicit waitsin 48 of 964, across 5 files
- Use bundled scripts as black boxesin 46 of 964, across 3 files
- Do not read script source codein 46 of 964, across 3 files
- Use sync playwright for scriptsin 46 of 964, across 3 files
- Inspect dom before executing actionsin 46 of 964, across 3 files
- Run the full test suitein 37 of 964
- Write the failing test firstin 29 of 964, across 23 files
Said here and by no other author read
- run the test repeatedly to reproduce the flakiness
- run tests in randomized order where relevant
- locate the source of non-determinism
- fix the root cause of non-determinism
- isolate test state with proper setup and teardown
- seed randomness and fake the clock
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.