agentsclimarketplace

Observability

Skill stevancris/sre-ai-agent/skills/observability

AI agent that accumulates SRE knowledge from every incident — built on Agent Skills spec for Claude Code

Install
npx -y skills add stevancris/sre-ai-agent --skill observability

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Design, review, and improve observability for services: metrics, logs, and distributed traces. Use when instrumenting a new service, investigating gaps in visibility, reviewing dashboards, building alerting, or diagnosing why an incident was hard to detect. Trigger keywords: observability, metrics, logs, tracing, traces, dashboard, alerting, instrumentation, OpenTelemetry, Prometheus, Datadog, Grafana, structured logging, spans, golden signals, USE method, RED method, four golden signals, visibility, monitoring gap, hard to debug, missing metrics, no alerts.

SKILL.md

5.4 KB, as published. Nobody here has run it

Observability Skill

Setup Check

Before loading context files, check if context/CONTEXT.md exists in the current directory.

If context/CONTEXT.md exists — read it and proceed normally.

If context/CONTEXT.md does not exist — this skill was installed standalone (e.g. via npx skills add). Ask the user these questions before proceeding:

  1. Rolejunior-sre / senior-sre / sre-manager (shapes output depth and tone)
  2. Cloud provideraws / gcp / azure / on-prem / hybrid
  3. Observability stack — e.g. Datadog, Prometheus+Grafana, New Relic
  4. Company name and primary services affected (if relevant to this task)

Use the answers inline for this session. For persistent setup across all skills, suggest:

pipx install sre-agent
sre-agent init

Instructions

Step 1: Load Context

Read context/company/tech-stack.md to identify the observability stack (Datadog, Prometheus+Grafana, New Relic, etc.) and primary language/framework.

Step 2: Determine the Mode

  • Instrument new service — adding observability to a new or existing service
  • Review coverage — assess gaps in existing observability
  • Build alerting — create alert rules for a service
  • Dashboard design — build or improve a service dashboard
  • Incident visibility gap — why was this incident hard to detect?

Mode: Instrument New Service

The Four Golden Signals (start here for every service)

SignalWhat to measureExample metric name
LatencyRequest duration, broken by success/errorhttp_request_duration_seconds
TrafficRequest rate (RPS / QPS)http_requests_total
ErrorsError rate (5xx, exceptions)http_errors_total
SaturationResource utilization (CPU, memory, queue depth)process_cpu_usage, queue_depth

Instrumentation Checklist

HTTP Services:

  • Request count (by method, path, status code)
  • Request latency histogram (P50, P95, P99)
  • Error rate (4xx and 5xx separately)
  • In-flight requests (saturation)
  • Dependency call latency (DB, cache, downstream APIs)
  • Circuit breaker state (if applicable)

Workers / Background Jobs:

  • Job execution count (success / failure)
  • Job duration histogram
  • Queue depth (consumer lag)
  • Job retry count
  • Dead-letter queue size

Databases:

  • Query latency (by query type)
  • Connection pool utilization
  • Replication lag (for read replicas)
  • Cache hit rate (if applicable)

Structured Logging Standards

Every log line must include:

{
  "timestamp": "ISO8601",
  "level": "INFO|WARN|ERROR",
  "service": "service-name",
  "version": "v1.2.3",
  "trace_id": "abc123",
  "span_id": "def456",
  "request_id": "uuid",
  "user_id": "optional",
  "message": "human readable",
  "error": "error string if applicable"
}

Never log: passwords, tokens, PII (email, phone, address), credit card numbers.

Distributed Tracing

  • Add trace context propagation at all service boundaries (HTTP headers, queue message attributes).
  • Create spans for: inbound requests, outbound calls, DB queries, cache operations.
  • Tag spans with: service name, operation name, status (ok/error), relevant business attributes.

Mode: Review Coverage

Assess against the four golden signals for each service in scope. Flag gaps:

Service: <name>
Golden Signal Coverage:
  Latency:    ✓ histogram available, P99 alert configured
  Traffic:    ✓ RPS metric available
  Errors:     ✗ MISSING — no error rate metric
  Saturation: ⚠ CPU only — memory and queue depth missing

Tracing:      ✓ spans present, but missing DB query spans
Logging:      ⚠ logs exist but unstructured (no trace_id)

Priority gaps:
  1. Add error rate metric (blocks SLO measurement)
  2. Add trace_id to log lines (needed for incident correlation)
  3. Add queue depth metric (saturation blind spot)

Mode: Build Alerting

For each service, build alert rules using the multi-window burn rate approach:

# Example: error rate alert (adapt syntax to your stack)
alert: HighErrorRate
expr: |
  (
    rate(http_errors_total[5m]) / rate(http_requests_total[5m])
  ) > 0.05
for: 2m
labels:
  severity: P1
annotations:
  summary: "Error rate above 5% for {{ $labels.service }}"
  runbook: "<link to runbook>"

Guidelines

  • Alert on symptoms (user-facing impact), not causes (CPU high).
  • Every alert must have a runbook. If no runbook exists, invoke runbook-generator.
  • Dashboards should tell a story from top (user experience) to bottom (infrastructure).
  • Do not alert on metrics that do not require human action.
  • Persona (junior-sre): add "learning note" explaining why each golden signal matters.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.