Resilience audit
Nine quality-canary skills for AI coding agents - code health, world-class rule completeness, grounding, supply chain, resilience, drift & more - plus the auto-cadence hooks that run them unprompted at session start and session end. Cross-agent, consent-gated, token-lean.
npx -y skills add HetCreep/CoalMine --skill resilience-auditAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 11 stars11 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Failure-mode audit (FMEA for software) — for each way the system can fail (network, storage, partial completion, crash, concurrency, bad input), check whether code DETECTS, HANDLES, RECOVERS, and COMMUNICATES it. Triggers on: "/resilience-audit", "resilience-audit", "FMEA audit". Use when touching network, storage, async, retry, or rollback paths. Flags data loss, silent-success-on-failure, missing rollback/retry/idempotency. Reports; does not fix unless asked.
SKILL.md
3.0 KB, as published. Nobody here has run it
Resilience Audit
<!-- SHARED:LANGUAGE_HEADER -->For every operation: "what happens when this FAILS?" Report; do NOT fix unless asked.
Failure categories
- External I/O — network down/slow, API 4xx/5xx/timeout, rate-limit. Retry w/ backoff? Timeout set? Clear error vs hang?
- Storage — disk full, permission denied, partial write. Atomic write (temp+rename)? Cleanup on failure? Existing good copy untouched?
- Partial completion — half-done op (50/100 files). Reported as FAILURE, never success.
- Crash / OOM — killed mid-op. Idempotent restart? No orphaned half-state?
- Concurrency — two instances, race, deadlock. Locking / idempotency / safe re-entry?
- Input / data — malformed, null, truncated, huge. Validate at boundary? Fail-fast?
- Dependency down — fallback/cache/graceful degrade? Clear error vs silent hang?
- Resource exhaustion — bounded? Backpressure? Cleanup on error path?
Per-stack timeout/atomicity/idempotency patterns to grep: read references/checks.md before scanning.
For each failure point, check 4 things
- Detected? code notices it (doesn't swallow)?
- Handled? retry/fallback/fail-clean — not ignored, not silent-success?
- Recoverable? rollback/idempotent; no data loss or corruption?
- Communicated? clear error to user+log; not a hang, not a false "done"?
Discipline
- Trace actual failure path (cite file:line). Don't assume handling exists; prove it.
- "partial = failure" — any path reporting success on partial completion = CRITICAL.
- "logged" ≠ "handled" — swallowed+logged error that corrupts state or returns success = CRITICAL.
Fix mode (choice-gated)
After the report, present via ask_question:
- Fix safe ones — add missing timeout, null/input validation, clear error+log on unhandled path. Each: checkpoint → fix → build+tests → revert if newly red.
- Let me pick — user-selected fixes only.
- Report only — change nothing.
NEVER auto-fix: retry/rollback/recovery/atomicity logic (semantic changes can introduce new failure modes).
Output
| operation | failure mode | effect | handling (file:line) | severity | recommended guard |
Ordering/atomicity findings · Summary (counts + top fixes) · Not assessed
Severity: CRITICAL (data loss/corruption/silent-success) · HIGH (crash/hang/partial-no-recovery) · MEDIUM (poor degradation/missing retry) · LOW (cosmetic)
<!-- SHARED:ORCHESTRATION --> <!-- SHARED:ESCALATION_FOOTER -->