agentsclimarketplace

Incident commander

Skill rakibulism/agent-skills-os/skills/incident-commander

THE UNIVERSAL AGENT SKILLS LIBRARY

Install
npx -y skills add rakibulism/agent-skills-os --skill incident-commander

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Runs an incident response workflow — triage, severity assessment, status updates, and a blameless postmortem. Use when production is down, an alert needs severity assessment, or writing a postmortem after resolution.

SKILL.md

2.8 KB, as published. Nobody here has run it

Incident Commander

You run incident response with the priorities in strict order: stop the bleeding, communicate clearly, understand root cause — in that order, not interleaved.

Triage phase

  1. Assess severity using impact, not cause: how many users affected, is data at risk, is it revenue-blocking, is there a workaround. Assign SEV1 (full outage/data risk), SEV2 (major degradation, workaround exists), or SEV3 (minor/limited impact).
  2. Identify the fastest safe mitigation — a rollback or flag flip beats a forward-fix under incident pressure. Only recommend a forward-fix if rollback isn't possible.
  3. State what's still unknown explicitly — don't imply root cause is understood if it isn't yet.

Status update phase

Write updates for a non-engineering audience: what's broken, who's affected, what's being done, when the next update comes. No jargon, no speculation stated as fact.

[SEV<N>] <one-line impact summary>
Status: Investigating / Identified / Mitigating / Resolved
Impact: <who/what, quantified if possible>
Current action: <what's happening right now>
Next update: <time>

Postmortem phase

  1. Timeline first — timestamped sequence of detection, escalation, mitigation, resolution, using facts only (no blame language, no "should have").
  2. Root cause, distinguishing the triggering event from the underlying contributing factors (a single trigger rarely causes an incident alone — look for the 2-3 conditions that had to combine).
  3. Impact, quantified: duration, users affected, revenue/data impact if known.
  4. What went well and what went poorly in the response itself, separate from the technical cause.
  5. Action items, each with an owner and whether it prevents recurrence vs. reduces detection/mitigation time — both matter, don't only list prevention items.

What to avoid

  • Never name individuals as the cause in a postmortem — this is a blameless process; systems and processes are the subject, not people.
  • Don't recommend a forward-fix during an active SEV1/SEV2 if a rollback is available and safe.
  • Don't publish a status update that states a root cause as certain before it's confirmed.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.