agentsclimarketplace

Incident response patterns

Skill vibeeval/vibecosystem/skills/incident-response-patterns

AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution.

Install
npx -y skills add vibeeval/vibecosystem --skill incident-response-patterns

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Incident severity classification, runbook templates, root cause analysis, and post-incident review

SKILL.md

2.8 KB, as published. Nobody here has run it

Incident Response Patterns

Severity Classification

LevelTanımResponse TimeEscalation
SEV-1Tam kesinti, data loss5 dkTüm ekip
SEV-2Major feature bozuk30 dkOn-call + lead
SEV-3Minor feature bozuk4 saatOn-call
SEV-4Kozmetik, low impactSonraki iş günüTicket

Incident Response Protocol

1. DETECT  → Alert tetiklendi / kullanıcı bildirdi
2. TRIAGE  → Severity belirle, incident channel aç
3. ASSIGN  → Incident Commander + Responder ata
4. CONTAIN → Bleeding'i durdur (rollback, feature flag off)
5. FIX     → Root cause'u düzelt
6. VERIFY  → Fix'i doğrula (monitoring, smoke test)
7. RESOLVE → Incident kapat, stakeholder bilgilendir
8. REVIEW  → Post-incident review (48 saat içinde)

Runbook Template

# Runbook: [Service Name] - [Issue Type]

## Symptoms
- Alert: [alert name]
- Dashboard: [link]
- Expected: [normal behavior]
- Actual: [observed behavior]

## Diagnosis Steps
1. Check [metric/log/dashboard]
2. Run: `kubectl get pods -n production`
3. Check: `SELECT count(*) FROM errors WHERE ...`

## Mitigation
1. Quick fix: [rollback/restart/scale]
2. Feature flag: [disable feature X]
3. Redirect: [failover to backup]

## Escalation
- L1: [on-call engineer]
- L2: [team lead]
- L3: [CTO/VP Eng]

Root Cause Analysis

5 Whys

Problem: API 500 hatası
Why 1: Database connection timeout
Why 2: Connection pool tükendi
Why 3: Slow query connection'ları tutuyor
Why 4: Missing index on frequently queried column
Why 5: Schema change review process eksik
→ Action: Schema change PR'da EXPLAIN ANALYZE zorunlu

Post-Incident Review Template

## Incident Summary
- Duration: [start - end]
- Impact: [affected users/revenue]
- Severity: SEV-[X]

## Timeline
- HH:MM - Alert triggered
- HH:MM - Incident declared
- HH:MM - Root cause identified
- HH:MM - Fix deployed
- HH:MM - Incident resolved

## Root Cause
[Detailed technical explanation]

## Action Items
- [ ] [Fix] Add missing index (owner: @dev, due: date)
- [ ] [Prevent] Add schema review step (owner: @lead)
- [ ] [Detect] Add connection pool alert (owner: @sre)

Checklist

  • Severity matrix tanımlı
  • On-call rotation aktif
  • Runbook'lar güncel
  • Incident channel template hazır
  • Post-incident review 48 saat içinde
  • Action items tracked (Jira/Linear)
  • Blameless culture enforced
  • Alerting threshold'lar doğru

Anti-Patterns

  • Blame culture (kişi değil, sistem düzelt)
  • Runbook'suz servis
  • Post-incident review atlamak
  • Alert fatigue (çok fazla false positive)
  • Hotfix without rollback plan

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.