agentsclimarketplace

Incident response

Skill stevancris/sre-ai-agent/skills/incident-response

Manage active production incidents end-to-end. Use when an alert fires, a service is degraded or down, error rates spike, latency increases, or a customer reports an outage. Covers detection, severity classification, response coordination, communication drafting, and incident timeline tracking. Trigger keywords: incident, outage, down, degraded, alert, pager, on-call, oncall, P0, P1, P2, SEV1, SEV2, firing, pagerduty, opsgenie, service unavailable, 5xx spike, error rate high, latency spike, customers affected.From its SKILL.md

Install
npx -y skills add stevancris/sre-ai-agent --skill incident-response

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

6.4 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it

Incident Response Skill

Setup Check

Before loading context files, check if context/CONTEXT.md exists in the current directory.

If context/CONTEXT.md exists — read it and proceed normally.

If context/CONTEXT.md does not exist — this skill was installed standalone (e.g. via npx skills add). Ask the user these questions before proceeding:

  1. Rolejunior-sre / senior-sre / sre-manager (shapes output depth and tone)
  2. Cloud provideraws / gcp / azure / on-prem / hybrid
  3. Observability stack — e.g. Datadog, Prometheus+Grafana, New Relic
  4. Company name and primary services affected (if relevant to this task)

Use the answers inline for this session. For persistent setup across all skills, suggest:

pipx install sre-agent
sre-agent init

Instructions

Step 0: Search Knowledge Base First (MANDATORY — before anything else)

Before classifying severity or opening a channel, search skills/knowledge-base/patterns/ using Glob (skills/knowledge-base/patterns/*.md) and Grep for the service name or symptom keywords from the user's message.

If a matching pattern file is found:

⚡ Known pattern detected: <pattern name>
Last seen: <date> | Times seen: <N>

Fastest diagnostic path:
  1. <step from pattern file>
  2. <step from pattern file>

What NOT to try (dead ends):
  - <dead end from pattern file>

Estimated resolution time based on history: ~X minutes

Surface this immediately — before opening the incident channel or doing anything else. If no pattern matches, continue to Step 1 and note: "No prior pattern found — starting fresh."

Step 1: Load Context

Read context/CONTEXT.md, context/company/incident-severity.md, and context/company/oncall-schedule.md before taking any other action.

Step 2: Classify Severity

Using the severity definitions from context/company/incident-severity.md, ask the user:

  • Is this user-facing?
  • What percentage of users are affected?
  • Which services or features are impacted?

Classify the incident as P0 / P1 / P2 / P3 based on the decision tree in the severity file.

Step 3: Open the Incident

Generate the following immediately:

Incident channel name:

inc-YYYY-MM-DD-<short-description>

Opening Slack message template:

🚨 INCIDENT [P<N>] — <short description>

Status: INVESTIGATING
Impact: <what users/services are affected>
Started: <approximate start time>
IC (Incident Commander): <primary oncall name>

Bridge: <link>
Runbook: <link if known>
Updates every 15 min or on status change.

Step 4: Immediate Actions Checklist

Output a checklist tailored to severity:

P0 checklist:

  • Notify primary on-call (if not already paged)
  • Notify secondary on-call
  • Notify engineering manager
  • Open incident bridge
  • Post status page update (draft below)
  • Identify blast radius: which services / regions / customers
  • Pull relevant runbook
  • Assign roles: IC (Incident Commander), Comms Lead, Tech Lead

P1 checklist:

  • Notify primary on-call (if not already paged)
  • Open incident Slack channel
  • Identify blast radius
  • Pull relevant runbook
  • Notify engineering manager

P2/P3 checklist:

  • Create incident ticket
  • Notify on-call team via Slack (no bridge needed)
  • Begin investigation

Step 5: Blast Radius Assessment

Determine and document:

  • Which services are affected (primary + downstream dependencies)?
  • Which regions are affected?
  • What percentage of users or requests are impacted?
  • Is there any data loss or data integrity risk?
  • Are any SLOs currently breaching?

Step 6: Pull the Runbook

Search skills/oncall-runbook/references/ for a runbook matching the affected service and alert name. If none exists, invoke the runbook-generator skill in parallel.

Step 7: Timeline Tracking

Maintain a running timeline as the user provides updates. Format:

HH:MM UTC — <event description>

Prompt the user for updates at natural pause points (after each diagnostic step).

Step 8: Status Updates

Generate stakeholder status updates every 15 minutes during P0/P1:

Update [N] — HH:MM UTC
Status: INVESTIGATING / IDENTIFIED / MITIGATING / RESOLVED
Impact: <current scope>
What we know: <current understanding>
Next update: HH:MM UTC

Step 9: Resolution

When the user signals the incident is resolved:

  1. Generate a final resolution message.
  2. Calculate MTTR (Mean Time To Resolve) from timeline.
  3. Prompt to transition to the postmortem skill.
  4. Remind to close the status page incident.

Examples

Example: Database connection pool exhausted

User: "Alert just fired — our API is returning 503s. Looks like DB connections exhausted."

Response (senior-sre persona):

Severity: P1 (API degraded, partial user impact)

Immediate actions:
- [ ] Open #inc-2024-01-15-db-conn-exhausted
- [ ] Check current connection count: kubectl exec <db-pod> -- psql -c "SELECT count(*) FROM pg_stat_activity;"
- [ ] Check which services are holding connections: kubectl logs -l app=api --since=5m | grep "connection"
- [ ] Scale API replicas down by 20% to shed load while diagnosing
- [ ] Check for connection leak in recent deploys: git log --oneline -10

Runbook: See skills/oncall-runbook/references/postgresql-connection-exhaustion.md

Guidelines

  • Never skip severity classification even if the user says "it's just a P2."
  • During a P0, produce the Slack message template before doing any technical diagnosis.
  • If the user does not know the blast radius, provide kubectl/query commands to find it.
  • Keep timeline entries timestamped to the minute; they become the postmortem foundation.
  • Persona: junior-sre version always ends each step with an escalation reminder.

What ships with it: 2 files

6.3 KB alongside SKILL.md

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.