agentsclimarketplace

Rca summary

Skill stevancris/sre-ai-agent/skills/rca-summary

AI agent that accumulates SRE knowledge from every incident — built on Agent Skills spec for Claude Code

Install
npx -y skills add stevancris/sre-ai-agent --skill rca-summary

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Generate a concise, shareable RCA (Root Cause Analysis) summary from raw incident data, postmortem notes, or a completed root-cause-analysis skill session. Use when you need to communicate findings quickly to stakeholders, produce a Slack or email digest of an RCA, create a one-pager for management, or add a brief RCA entry to a knowledge base. Distinct from the full postmortem or deep root-cause-analysis skill — this produces a tight, readable summary optimized for sharing. Trigger keywords: RCA summary, summarize RCA, RCA one-pager, share the RCA, brief RCA, quick summary of what happened, executive RCA, RCA for management, send the findings, incident summary, what happened summary, TL;DR of the incident, incident digest, share postmortem findings, RCA writeup, brief postmortem.

SKILL.md

8.2 KB, as published. Nobody here has run it

RCA Summary Skill

Setup Check

Before loading context files, check if context/CONTEXT.md exists in the current directory.

If context/CONTEXT.md exists — read it and proceed normally.

If context/CONTEXT.md does not exist — this skill was installed standalone (e.g. via npx skills add). Ask the user these questions before proceeding:

  1. Rolejunior-sre / senior-sre / sre-manager (shapes output depth and tone)
  2. Cloud provideraws / gcp / azure / on-prem / hybrid
  3. Observability stack — e.g. Datadog, Prometheus+Grafana, New Relic
  4. Company name and primary services affected (if relevant to this task)

Use the answers inline for this session. For persistent setup across all skills, suggest:

pipx install sre-agent
sre-agent init

Instructions

Step 1: Load Context

Read context/CONTEXT.md for persona. The output format varies significantly by persona:

  • junior-sre: learning-focused, detailed enough to build mental model
  • senior-sre: dense technical digest, assumes shared context with the team
  • sre-manager: business-impact-first, ready to forward to executives

Step 2: Gather Source Material

Accept input from any of these sources (ask the user which applies):

  • A completed postmortem skill session in this conversation
  • A completed root-cause-analysis skill session in this conversation
  • Raw text pasted by the user (incident channel log, notes, timeline)
  • A postmortem document file path (read it with the Read tool)

Step 3: Extract the Key Facts

From the source material, extract:

FieldSource
Incident name / IDpostmortem title or incident channel name
Date and time (UTC)incident timeline
DurationRESOLVED time − STARTED time
SeverityP0 / P1 / P2 / P3
Services affectedblast radius section
User impactimpact section
What happenedsymptom description
Why it happenedroot cause from five-whys
How it was fixedresolution section
How it was detectedtimeline (DETECTED entry)
TTD / TTM / MTTRcalculated from timeline
Top action itemsaction items table

Step 4: Generate the RCA Summary

Produce the summary in the format matching the requested audience:


Format A: Slack / Teams Digest (default)

Optimized for posting in an incident channel or engineering Slack channel. Aim for < 300 words, scannable with bold headings.

*RCA Summary — [Incident Name]* | [Date] | [Duration] | [Severity]

*What happened*
[2–3 sentences. What did users experience? When did it start and end?]

*Root cause*
[1–2 sentences. The systemic reason this was possible, not just the proximate cause.]

*Contributing factors*
• [Factor 1]
• [Factor 2]

*How we detected it*
[1 sentence. Alert, user report, or manual discovery. How long after the problem started?]

*How we fixed it*
[1–2 sentences. Immediate remediation taken.]

*Key metrics*
• Time to detect: Xm | Time to mitigate: Xm | MTTR: Xh Ym
• Users affected: ~N (X%)

*Top action items*
1. [Action] — [Owner role] — [Due date]
2. [Action] — [Owner role] — [Due date]
3. [Action] — [Owner role] — [Due date]

Full postmortem: [link or "in progress"]

Format B: Executive / Management Email

Optimized for forwarding to VP, CTO, or customer success. Non-technical language. Aim for < 200 words.

Subject: Incident Summary — [Service/Feature], [Date]

On [date], [service or feature] experienced [plain-language description of user impact]
for approximately [duration]. This affected approximately [N] users ([X]% of our user base).

Root cause: [one sentence in plain language — no jargon]

We resolved the incident by [resolution description].

To prevent recurrence, we are:
1. [Action item 1 — outcome-focused, not technical]
2. [Action item 2]
3. [Action item 3]

A full technical postmortem is [in progress / available here: link].

Please let me know if you have any questions.

Format C: Knowledge Base / Confluence Entry

Optimized for long-term reference. Structured for searchability.

# RCA: [Incident Name]
**Date:** YYYY-MM-DD | **Severity:** P<N> | **Duration:** Xh Ym | **MTTR:** Xh Ym

## What Happened
[2–3 sentences. Observable impact from user perspective.]

## Timeline (condensed)
- `HH:MM` Started
- `HH:MM` Detected (TTD: Xm)
- `HH:MM` Mitigated (TTM: Xm)
- `HH:MM` Resolved (MTTR: Xm)

## Root Cause
[1–2 sentences. Systemic cause.]

## Proximate Cause
[1 sentence. Immediate technical trigger.]

## Contributing Factors
- **Detection gap:** [why it took Xm to detect]
- **Response gap:** [what slowed response, if applicable]
- **Process gap:** [what process was missing]

## Resolution
[What was done to stop the impact.]

## Action Items
| # | Action | Owner | Due | Status |
|---|---|---|---|---|
| 1 | | | | Open |
| 2 | | | | Open |

## Tags
`[service-name]` `[failure-mode]` `[team-name]`

Full postmortem: [link]

Format D: One-Line Digest (for incident retrospective lists)

For weekly/monthly incident digest tables:

| YYYY-MM-DD | P<N> | [Service] | [Xh Ym] | [Root cause in 10 words] | [Top action item] |

Step 5: Offer Format Selection

If the user did not specify a format, ask: "Which format would you like? A) Slack digest (default) B) Executive email C) Knowledge base / Confluence entry D) One-line digest"

Or produce Format A automatically and offer to convert to others.

Step 6: Follow-Up Prompts

After generating the summary, offer:

  • "Want me to also generate a one-line version for your monthly incident digest?"
  • "Should I add this to the knowledge base (Format C)?"
  • "Would you like a version tailored for customer communication?"

Examples

Example Slack digest (senior-sre persona)

*RCA Summary — DB Connection Pool Exhaustion* | 2024-01-15 | 47 min | P1

*What happened*
Checkout failures (503s) for ~15% of users between 14:03–14:50 UTC.
Impact: estimated 2,400 failed checkout attempts.

*Root cause*
No automated load testing in CI allowed an N+1 query regression (introduced in v2.4.0)
to reach production undetected.

*Contributing factors*
• No connection pool saturation alert (detection took 42 min post-deploy)
• Staging does not mirror production traffic patterns

*How we detected it*
PagerDuty latency P99 alert fired at 14:03 UTC — 42 minutes after deploy completed.

*How we fixed it*
Rolled back to v2.3.9 at 14:45 UTC. Restored in 7 minutes.

*Key metrics*
• TTD: 42m | TTM: 7m | MTTR: 47m
• Users affected: ~2,400 (~15% of active users in window)

*Top action items*
1. Add DB connection pool saturation alert — SRE team — Jan 22
2. Add load test to CI pipeline — Backend team — Feb 5
3. Implement canary deploy for all services — Platform team — Feb 20

Full postmortem: [link]

Guidelines

  • The root cause in a summary must be the systemic cause, not "the deploy broke it."
  • Never include names of individuals in the summary — only teams and roles.
  • TTD / TTM / MTTR must be calculated accurately from the timeline; do not estimate.
  • Persona (sre-manager): always lead with user/business impact before technical detail, and ensure action items are outcome-focused ("reduce MTTR by X%" not "add an alert").
  • If source material is incomplete (missing timeline, no root cause identified yet): flag the gaps clearly and produce a partial summary with [TBD] placeholders. Do not fabricate data.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.