agentsclimarketplace

Postmortem

Skill stevancris/sre-ai-agent/skills/postmortem

AI agent that accumulates SRE knowledge from every incident — built on Agent Skills spec for Claude Code

Install
npx -y skills add stevancris/sre-ai-agent --skill postmortem

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Create blameless postmortems after incidents are resolved. Use when an incident is over, during post-incident review, or when writing a retrospective. Covers timeline reconstruction, contributing factor analysis, action item generation, and distribution. Trigger keywords: postmortem, post-mortem, post incident review, PIR, retrospective, blameless, lessons learned, what went wrong, action items, incident review, follow-up, write up, incident report.

SKILL.md

6.1 KB, as published. Nobody here has run it

Postmortem Skill

Setup Check

Before loading context files, check if context/CONTEXT.md exists in the current directory.

If context/CONTEXT.md exists — read it and proceed normally.

If context/CONTEXT.md does not exist — this skill was installed standalone (e.g. via npx skills add). Ask the user these questions before proceeding:

  1. Rolejunior-sre / senior-sre / sre-manager (shapes output depth and tone)
  2. Cloud provideraws / gcp / azure / on-prem / hybrid
  3. Observability stack — e.g. Datadog, Prometheus+Grafana, New Relic
  4. Company name and primary services affected (if relevant to this task)

Use the answers inline for this session. For persistent setup across all skills, suggest:

pipx install sre-agent
sre-agent init

Instructions

Step 1: Load Context and Incident Data

Read context/CONTEXT.md for persona. Ask the user for:

  • Incident name / ID
  • Incident timeline (or pull from incident-response skill if available)
  • Severity classification
  • Affected services and estimated user impact
  • Resolution summary

Step 2: Reconstruct the Timeline

Build a complete chronological timeline with these milestone markers:

  • [STARTED] — when did the problem actually begin (may predate detection)
  • [DETECTED] — when did monitoring or a user report alert the team
  • [ACKNOWLEDGED] — when did someone start working on it
  • [IDENTIFIED] — when was the root cause identified
  • [MITIGATED] — when was user impact stopped (may differ from root cause fix)
  • [RESOLVED] — when was the system fully restored

Calculate and highlight:

  • Time to Detect (TTD): DETECTED − STARTED
  • Time to Acknowledge (TTA): ACKNOWLEDGED − DETECTED
  • Time to Mitigate (TTM): MITIGATED − ACKNOWLEDGED
  • MTTR: RESOLVED − STARTED

Step 3: Root Cause Analysis

Apply the five-whys technique. Start from the customer-facing symptom and ask "why" iteratively until reaching a systemic or process-level cause.

Format:

Why 1: Why did users see 503 errors?
  → Because the API pods were crashing.
Why 2: Why were the API pods crashing?
  → Because they ran out of memory.
Why 3: Why did they run out of memory?
  → Because a memory leak was introduced in the v2.3.1 deploy.
Why 4: Why was the memory leak not caught before deploy?
  → Because the staging environment does not run load tests.
Why 5: Why does staging not run load tests?
  → Because we have no automated load testing in CI.
Root Cause: No automated load testing in CI allowed a memory leak to reach production.

Step 4: Contributing Factor Analysis

Categorize contributing factors across four dimensions:

  • Detection gaps — why was it hard to detect or took too long?
  • Response gaps — why did response take longer than it should have?
  • Process gaps — what process or procedure failed or was missing?
  • Tooling gaps — what tooling limitation made this worse?

Step 5: Generate Action Items

For each gap identified, generate a SMART action item:

  • What: specific change to make
  • Why: which gap it closes
  • Owner: team or role responsible (do not name individuals; assign to roles)
  • Due: relative date (e.g., "2 weeks", "next sprint")
  • Priority: P0 (blocks similar incident) / P1 (reduces risk significantly) / P2 (improvement)

Step 6: Produce the Postmortem Document

Fill in the template from references/postmortem-template.md.

Persona adjustments:

  • sre-manager: prepend an executive summary with: total downtime, estimated revenue impact (if known), number of users affected, and top 1–2 action items.
  • junior-sre: add a "What I learned" section at the end with educational notes.

Step 7: Distribution

Recommend the distribution list based on severity:

  • P0: all engineering, customer success, executive team, shared to public status page
  • P1: engineering team, engineering manager, customer success
  • P2/P3: SRE team and service owners

Step 8: Close the Knowledge Loop (MANDATORY)

After the postmortem is written and distributed, always prompt:

"Postmortem complete. Before we close — run knowledge-capture to save this as a searchable pattern. Next time this happens, any SRE on the team can resolve it faster."

This step converts the postmortem (written for humans, looking backward) into a knowledge-base pattern (written for the agent, looking forward). The two documents serve different purposes and both are needed.

If the user already ran knowledge-capture during the incident, confirm the pattern exists in skills/knowledge-base/patterns/ and prompt them to enrich it with any additional detail from the postmortem (dead ends, prevention action items, MTTR).


Examples

Example postmortem action items

PriorityActionOwnerDue
P0Add memory usage alert at 80% thresholdSRE team1 week
P0Add load test to staging CI pipelineBackend team2 weeks
P1Add canary deployment step for all servicesPlatform team1 month
P2Document memory leak debugging runbookSRE team2 weeks

Guidelines

  • Postmortems are blameless: do not name individuals, only systems, processes, and teams.
  • The root cause should always be a systemic issue, never "human error." ("Human error" is always a symptom — the root cause is why the error was possible.)
  • Action items without owners and due dates will not get done; enforce SMART format.
  • The five-whys should reach a process or systemic level by Why 4 or 5.
  • Never close a postmortem with "we will be more careful next time."

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.