Incident runbook
Skill byerlikaya/claude-starter-kit/claude-starter/skills/incident-runbook
Production incident response: diagnose → mitigate → resolve, then a blameless postmortem and a repeatable runbook. Stop the impact first, root cause second. Trigger phrases: "incident", "incident response", "runbook", "postmortem", "root cause", "outage", "production incident", "post-incident"From its SKILL.md
npx -y skills add byerlikaya/claude-starter-kit --skill incident-runbookAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- 22 stars22 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
- runs commandsInstructs the agent to run 2 commands, including `vps-deploy revert` and 1 more.
SKILL.md
2.0 KB, 412 tokens by cl100k_base, as published. Nobody here has run it
Incident Response & Runbook
Two modes: live incident (what to do right now) and aftermath (postmortem + runbook). Priority: stopping user impact > finding the root cause. No panic, one ordered step at a time.
Live incident — sequence
- Acknowledge & classify — what is the impact (who, how much), severity (SEV1 full outage … SEV3 minor).
- Mitigate the impact FIRST — rollback, turn off a feature flag, shift traffic, scale up. Without waiting on the root cause.
- Single coordinator — it is clear who decides; communication goes through one channel.
- Diagnose — last change? (deploy/migration/config) narrow it down with logs+metrics+traces (observability).
- Resolve — the smallest safe fix; then verify (health check).
- Close — confirm the impact is over; note the timeline (a postmortem input).
Mitigation reflexes
- Last deploy suspect → rollback (vps-deploy revert).
- Suspect feature → turn off the feature flag.
- After a destructive migration → restore from backup (db-migration).
- Dependency/service down → circuit breaker / graceful degradation.
After the incident
Blameless postmortem + producing a durable runbook: references/postmortem.md.
Invariant rules
- Stop the impact, then understand — the root cause does not hold up the resolution.
- Blameless culture — the postmortem questions the system, not the person.
- Actions are owned + dated — no "we'll look at it later".
- The runbook is executable — real commands/steps, not wishes.
- Make learning permanent — the lesson goes into an adr/runbook/monitoring, it does not get lost.
What ships with it: 1 file
885 B alongside SKILL.md
references/
- postmortem.md885 B