agentsclimarketplace

Incident response

Skill MARUCIE/openclaw-foundry/web/public/packs/spellbook-platform-engineer/skills/incident-response

The curated AI Agent skill marketplace — 37K+ vetted skills, S/A/B/C ratings, deploy anywhere

Install
npx -y skills add MARUCIE/openclaw-foundry --skill incident-response

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when triaging a production alert, writing a postmortem, creating or updating a runbook, classifying incident severity, or setting up on-call escalation paths.

SKILL.md

12.7 KB, as published. Nobody here has run it

是什么

这是一份事故响应规范,覆盖告警分级、值班轮换、应急处置、复盘流程,让团队遇到生产事故时不再手忙脚乱,而是按既定 runbook(应急手册)分钟级介入,事后还能沉淀成可复用的预防机制。

怎么用

  1. 接到告警时,先按严重度分级表(P0/P1/P2/P3)判断响应等级,确定要不要拉群、要不要通知客户。
  2. 处置过程套用文档里的诊断模板,先稳定再修复,避免越查越乱。
  3. 同类问题第二次出现时,立刻把处置步骤沉淀成 runbook,下次同事看一眼就能处理。
  4. 事故关闭后 48 小时内按本文档的复盘模板写 postmortem(事故复盘),重点找系统性根因不是怪个人。
  5. 每月统计 MTTR(平均恢复时长)和复发率两个指标,超阈值就组织专项治理。

架构图

flowchart LR
    A[告警触发] --> B[分级判断]
    B --> C[值班介入]
    C --> D[按 Runbook 处置]
    D --> E[业务恢复]
    E --> F[复盘沉淀]

Incident Response

Incident response is the structured process of detecting, mitigating, communicating, and learning from production failures to minimise user impact and prevent recurrence.

When to Activate

  • Triaging a production alert or on-call page
  • Writing a postmortem after an incident
  • Creating or updating a runbook for a service
  • Defining severity levels and escalation paths for a team
  • Setting up an on-call rotation
  • Running an incident response drill or game day

Severity Classification

SeverityDefinitionResponse SLAComms cadenceExample
P0Total outage or data loss — all users affectedPage immediately, < 5 minEvery 15 minPayment service down, DB unreachable
P1Major feature broken — most users affected< 15 min acknowledgementEvery 30 minLogin failing for 50%+ of users
P2Significant degradation — subset of users affected< 1 hourEvery 2 hoursSearch slow for US region
P3Minor issue — small impact, workaround availableNext business dayOnce resolvedNon-critical dashboard shows stale data
P4Cosmetic / no user impactSprint backlogN/ALog noise, minor UI misalignment

Escalation path:

  • P0/P1: page on-call engineer → page on-call lead if not ack'd in 5 min → escalate to eng manager
  • P2: page on-call engineer
  • P3/P4: create ticket, no page

Incident Lifecycle

Detection → Triage → Mitigate → Communicate → Resolve → Review (Postmortem)

First 5 Minutes — Triage Checklist

  • Acknowledge the alert and claim the incident in your incident tool (PagerDuty / Opsgenie)
  • Identify: what is broken, who is affected, since when?
  • Check the deployment timeline: was anything deployed in the last 2 hours?
  • Check the dashboards: error rate, latency, saturation — which service is the origin?
  • Open an incident channel: #inc-YYYY-MM-DD-short-description
  • Post initial acknowledgement message (see template below)
  • Assign roles: Incident Commander (IC), Communicator, Subject Matter Expert (SME)

Communication Templates

Initial Acknowledgement

🔴 [P0/P1 INCIDENT] Payment service degradation

Status: Investigating
Impact: ~30% of payment requests failing with 500 errors since 14:23 UTC
Affected: All users attempting checkout

IC: @alice
SME: @bob
Next update: 14:45 UTC

Tracking: https://incident.example.com/inc-2024-0042

Status Update (every 15–30 min for P0/P1)

🟡 [P1 UPDATE] Payment service — 14:45 UTC

Status: Mitigating
Root cause identified: Connection pool exhaustion after deploy at 14:15
Action taken: Rolled back to v2.3.1, monitoring error rate
Current error rate: 2% (down from 30%)

Next update: 15:00 UTC

Resolution

✅ [P1 RESOLVED] Payment service — 15:02 UTC

Status: Resolved
Duration: 39 minutes (14:23 – 15:02 UTC)
Root cause: Deploy v2.4.0 introduced a connection leak; pool exhausted under load
Resolution: Rolled back to v2.3.1; error rate returned to baseline at 15:00

Users impacted: ~15,000 failed checkout attempts
Follow-up: Postmortem scheduled for 2024-01-16 15:00 UTC
Incident report: https://incident.example.com/inc-2024-0042

Mitigation Decision Tree

Error rate > SLO threshold?
├── Yes
│   ├── Was something deployed in the last 2 hours?
│   │   ├── Yes → ROLLBACK first, investigate after
│   │   └── No  → Check: DB, cache, upstream dependency, config change
│   ├── Can we isolate the impact with a feature flag kill?
│   │   └── Yes → Kill the flag immediately
│   └── Is this a traffic spike?
│       └── Yes → Scale up horizontally, enable circuit breaker
└── No — latency degraded only?
    ├── Check DB: slow queries, lock contention, pool saturation
    ├── Check cache hit rate: has cache been evicted?
    └── Check upstream service latency

When NOT to roll back immediately:

  • The new version fixes a critical security issue (rolling back re-introduces the vulnerability)
  • Rollback would itself cause data migration issues
  • The issue is cosmetic (P3/P4) and the fix is already in progress

Runbook Structure

Runbooks must be written for the 3am engineer who has never seen this service.

# Runbook: [Service Name] — [Alert Name]

## Service Overview
[2–3 sentences: what does this service do, what does it depend on?]

## Alert: [Alert Name]
**Trigger condition:** [e.g., error rate > 1% for 5 minutes]
**Severity:** P1
**Dashboard:** [link]
**Logs:** [link to log query]

## Diagnostic Steps
1. Check the error rate panel on the [service dashboard](link)
   - Expected: < 0.1%
   - If > 1%: proceed to step 2
2. Check recent deployments:
   ```bash
   kubectl rollout history deployment/payment-service -n production
  1. Check DB connection pool:
    kubectl exec -it $(kubectl get pod -l app=payment-service -o name | head -1) \
      -- curl -s localhost:8080/metrics | grep db_pool
    
    • If db_pool_wait_duration_seconds > 1s: pool is exhausted, proceed to step 4
  2. Check for slow queries:
    SELECT query, mean_exec_time, calls
    FROM pg_stat_statements
    ORDER BY mean_exec_time DESC
    LIMIT 10;
    

Mitigation Steps

  • If recent deployment: kubectl rollout undo deployment/payment-service -n production
  • If DB pool exhausted: Scale up replicas: kubectl scale deployment/payment-service --replicas=6
  • If upstream dependency: Enable circuit breaker feature flag: [link to flag]

Escalation

  • If not resolved in 30 minutes: page @payment-team-lead
  • DB issues: page @dba-on-call
  • Infrastructure: page @infra-on-call

Related Runbooks


**Runbook quality checks:**
- Every step has an expected output — the engineer knows what "normal" looks like
- Commands are copy-paste ready (no placeholders that need substitution)
- Decision points have explicit branches ("if X, do Y; if Z, do W")
- Links to dashboards, log queries, and escalation contacts are current

## Blameless Postmortem

Write the postmortem within 48 hours while details are fresh. **Blameless = focus on systems and processes, not individuals.**

```markdown
# Postmortem: [Service] [Brief Description] — [Date]

## Summary
[2–3 sentences: what happened, impact, how it was resolved]

**Impact:** [number of users affected, % error rate, duration]
**Detection time:** [how long from start to detection]
**Resolution time:** [how long from detection to resolution]

## Timeline (UTC)
| Time  | Event |
|-------|-------|
| 14:15 | Deploy v2.4.0 rolled out to 100% |
| 14:23 | Alert fired: error rate > 1% |
| 14:28 | On-call acknowledged, started investigation |
| 14:38 | Root cause identified: connection pool exhausted |
| 14:45 | Rollback initiated |
| 15:00 | Error rate returned to baseline |
| 15:02 | Incident declared resolved |

## Root Cause Analysis (5 Whys)
1. **Why** did payment requests fail?
   → DB connection pool was exhausted
2. **Why** was the pool exhausted?
   → v2.4.0 introduced a connection leak in the retry handler
3. **Why** did the retry handler leak connections?
   → The `defer conn.Close()` was placed inside the retry loop, closing on each attempt but not releasing the acquired connection back to the pool
4. **Why** wasn't this caught in testing?
   → Integration tests used a single-connection test DB; pool exhaustion only manifests at scale
5. **Why** wasn't this caught by the integration test DB pool?
   → Test pool size was set to 100 (no practical limit); prod pool size is 20

## Contributing Factors
- No load test run before this deploy
- No DB pool exhaustion alert existed
- Code review missed the subtle connection lifecycle issue

## What Went Well
- Alert fired within 8 minutes of degradation starting
- On-call was paged and acknowledged quickly
- Rollback decision was made in < 10 minutes

## Action Items

| Action | Owner | Due | Category |
|--------|-------|-----|----------|
| Add DB pool wait time alert (threshold: > 1s for 5 min) | @alice | 2024-01-19 | Detection |
| Add integration test that simulates pool exhaustion under concurrent load | @bob | 2024-01-26 | Prevention |
| Add `db_pool_size` check to pre-deploy checklist | @alice | 2024-01-19 | Prevention |
| Run k6 load test before all deploys touching DB connection code | @bob | 2024-01-26 | Prevention |

Action Item Categories

  • Prevention: stops this class of failure from happening
  • Detection: reduces time-to-detection (MTTD)
  • Response: reduces time-to-resolution (MTTR)

Metrics to Track

MetricDefinitionTarget
MTTDMean Time To Detect — start of incident to first alert firing< 5 min
MTTAMean Time To Acknowledge — alert fires to on-call acks< 5 min
MTTRMean Time To Resolve — detection to resolution< 30 min for P0/P1
Incident frequencyNumber of P0/P1 incidents per month per serviceTrack trend; goal: decreasing
Repeat incidentsIncidents with the same root cause as a prior incidentGoal: 0

Review these monthly per service. Rising MTTR = runbooks need updating. Repeat incidents = action items not implemented.

See also: observability, deployment-strategies

Red Flags

  • Postmortem that names individuals as root cause — "Alice deployed bad code" stops at the human rather than the system that allowed the bad code to reach production; blameless postmortems ask why the system made it possible
  • Action items with no owner or no due date — "Improve monitoring" as an action item is never done; every item must have a named owner and a specific due date to be tracked and closed
  • Runbook that assumes the on-call engineer knows the service — runbooks must include what "normal" looks like and copy-paste commands; a 3am engineer touching an unfamiliar service cannot safely improvise
  • Rolling back immediately without checking if the rollback itself causes data loss — rolling back a deploy that ran a destructive migration may orphan or corrupt rows that were written against the new schema
  • Posting a P0 incident only in an engineering Slack channel — stakeholders (product, support, leadership) need timely updates via their own channels; the Communicator role exists specifically to bridge this gap
  • Severity P0 declared for every outage regardless of blast radius — "P0" becomes meaningless if used for single-user bugs; a calibrated P0 ensures the right resources are mobilized and avoids on-call fatigue
  • MTTD and MTTR tracked per-incident but never aggregated — individual numbers without a monthly trend hide whether the team is improving; review rolling averages per service each month
  • Closing an incident before a postmortem is scheduled — if the postmortem is not scheduled at resolution time it rarely happens; require a postmortem date as a condition of closing any P0 or P1

Checklist

  • Incident acknowledged within SLA (P0: 5 min, P1: 15 min)
  • Incident channel opened and IC/SME roles assigned
  • Initial acknowledgement posted to stakeholder channel
  • Status updates sent on cadence (every 15 min for P0, 30 min for P1)
  • Resolution announcement sent with impact summary
  • Postmortem written within 48 hours of resolution
  • 5 Whys root cause analysis complete (not just "human error")
  • Action items are SMART: owner, due date, and category (prevention/detection/response)
  • Runbook updated based on lessons learned
  • MTTD, MTTA, MTTR recorded for this incident

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.