Observability engineering
The definitive collection of cross-platform Agent Skills. Compatible with Claude Code, Codex, Cursor, OpenClaw, Gemini CLI, Copilot, Hermes. Curated weekly. Higher quality than any alternative.
npx -y skills add JPeetz/agent-skills --skill observability-engineeringAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Production-grade observability engineering for AI agents: OpenTelemetry instrumentation, monitoring setup, log aggregation, distributed tracing, SLI/SLO management, alert design, and incident response workflows. Implements vendor-neutral telemetry standards with structured incident management.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
37.0 KB, as published. Nobody here has run it
Observability Engineering
Production-grade observability engineering for AI agents. Covers the full observability lifecycle: OpenTelemetry instrumentation, metrics collection, structured logging, distributed tracing, SLI/SLO management, alert design, and incident response workflows.
When to Use This Skill
Invoke this skill when the user asks to:
- Instrument a service, application, or library with OpenTelemetry
- Set up monitoring dashboards, alerts, or metrics pipelines (Prometheus, Grafana, Datadog)
- Design SLOs/SLIs with error budgets and burn-rate alerts
- Configure distributed tracing with sampling strategies and context propagation
- Aggregate logs with structured JSON logging, trace correlation, and PII redaction
- Build incident response runbooks, communication templates, and postmortems
- Manage observability-as-code via Terraform/Pulumi for dashboards and alerts
- Optimize observability costs through cardinality management and retention policies
Do NOT use this skill for: general bug fixes (use code-review), Kubernetes deployment configuration (use a k8s skill), or generic DevOps questions without an observability intent.
1. OpenTelemetry Instrumentation
1.1 Quick-Start Patterns by Language
Node.js / TypeScript
// packages: @opentelemetry/api @opentelemetry/sdk-node @opentelemetry/auto-instrumentations-node
// @opentelemetry/exporter-trace-otlp-http @opentelemetry/exporter-metrics-otlp-http
// @opentelemetry/sdk-logs @opentelemetry/exporter-logs-otlp-http
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { OTLPMetricExporter } from '@opentelemetry/exporter-metrics-otlp-http';
import { PeriodicExportingMetricReader } from '@opentelemetry/sdk-metrics';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: `${process.env.OTEL_EXPORTER_OTLP_ENDPOINT}/v1/traces`,
}),
metricReader: new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({
url: `${process.env.OTEL_EXPORTER_OTLP_ENDPOINT}/v1/metrics`,
}),
exportIntervalMillis: 15000,
}),
instrumentations: [getNodeAutoInstrumentations()],
serviceName: process.env.OTEL_SERVICE_NAME || 'my-service',
});
sdk.start();
process.on('SIGTERM', () => sdk.shutdown().then(() => process.exit(0)));
Python
# packages: opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp
# opentelemetry-instrumentation-flask opentelemetry-instrumentation-requests
from opentelemetry import trace, metrics
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.resources import SERVICE_NAME, Resource
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
resource = Resource(attributes={SERVICE_NAME: "my-service"})
# Traces
provider = TracerProvider(resource=resource)
processor = BatchSpanProcessor(OTLPSpanExporter())
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)
# Metrics
metric_reader = PeriodicExportingMetricReader(OTLPMetricExporter())
meter_provider = MeterProvider(resource=resource, metric_readers=[metric_reader])
metrics.set_meter_provider(meter_provider)
Go
// modules: go.opentelemetry.io/otel go.opentelemetry.io/otel/sdk
// go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp
// go.opentelemetry.io/otel/exporters/otlp/otlpmetric/otlpmetrichttp
import (
"context"
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
sdkmetric "go.opentelemetry.io/otel/sdk/metric"
semconv "go.opentelemetry.io/otel/semconv/v1.26.0"
)
func initOtel(ctx context.Context) (*sdktrace.TracerProvider, *sdkmetric.MeterProvider, error) {
res, _ := resource.New(ctx,
resource.WithAttributes(semconv.ServiceName("my-service")),
)
tp := sdktrace.NewTracerProvider(
sdktrace.WithResource(res),
sdktrace.WithBatcher(otlptracehttp.New(ctx)),
)
otel.SetTracerProvider(tp)
mp := sdkmetric.NewMeterProvider(
sdkmetric.WithResource(res),
sdkmetric.WithReader(otlpmetrichttp.NewReader(ctx)),
)
otel.SetMeterProvider(mp)
return tp, mp, nil
}
Java
// dependencies: opentelemetry-bom, opentelemetry-exporter-otlp
// Run with: java -javaagent:opentelemetry-javaagent.jar -jar app.jar
// Auto-instrumentation is the recommended approach for Java.
// Manual configuration (Spring Boot example):
@Configuration
public class OpenTelemetryConfig {
@Bean
public OpenTelemetry openTelemetry() {
Resource resource = Resource.getDefault()
.merge(Resource.create(Attributes.of(
ResourceAttributes.SERVICE_NAME, "my-service")));
SdkTracerProvider tracerProvider = SdkTracerProvider.builder()
.addSpanProcessor(BatchSpanProcessor.builder(
OtlpHttpSpanExporter.builder().build()).build())
.setResource(resource)
.build();
SdkMeterProvider meterProvider = SdkMeterProvider.builder()
.registerMetricReader(PeriodicMetricReader.builder(
OtlpHttpMetricExporter.builder().build()).build())
.setResource(resource)
.build();
return OpenTelemetrySdk.builder()
.setTracerProvider(tracerProvider)
.setMeterProvider(meterProvider)
.build();
}
}
.NET
// packages: OpenTelemetry, OpenTelemetry.Exporter.OpenTelemetryProtocol
using OpenTelemetry;
using OpenTelemetry.Resources;
using OpenTelemetry.Trace;
using OpenTelemetry.Metrics;
var resourceBuilder = ResourceBuilder.CreateDefault()
.AddService("my-service");
using var tracerProvider = Sdk.CreateTracerProviderBuilder()
.SetResourceBuilder(resourceBuilder)
.AddOtlpExporter()
.AddAspNetCoreInstrumentation()
.AddHttpClientInstrumentation()
.Build();
using var meterProvider = Sdk.CreateMeterProviderBuilder()
.SetResourceBuilder(resourceBuilder)
.AddOtlpExporter()
.AddAspNetCoreInstrumentation()
.AddRuntimeInstrumentation()
.Build();
Ruby
# gems: opentelemetry-sdk opentelemetry-exporter-otlp
# opentelemetry-instrumentation-all
require 'opentelemetry/sdk'
require 'opentelemetry/exporter/otlp'
OpenTelemetry::SDK.configure do |c|
c.service_name = 'my-service'
c.use_all # auto-instrument all registered libraries
c.add_span_processor(
OpenTelemetry::SDK::Trace::Export::BatchSpanProcessor.new(
OpenTelemetry::Exporter::OTLP::Exporter.new
)
)
end
1.2 Auto-Instrumentation vs Manual Instrumentation
| Approach | When to Use | Pros | Cons |
|---|---|---|---|
| Auto-instrumentation | HTTP frameworks, DB clients, gRPC, messaging | Zero code changes, fast coverage | Less semantic depth, some noise |
| Manual spans | Business logic, custom operations, critical paths | Full semantic control, business context | Requires code changes, risk of gaps |
| Hybrid (recommended) | Production services | Best coverage + business context | Requires planning |
Auto-instrumentation agents:
| Language | Agent/Approach |
|---|---|
| Node.js | @opentelemetry/auto-instrumentations-node or --require @opentelemetry/auto-instrumentations-node/register |
| Python | opentelemetry-instrument CLI wrapper |
| Java | opentelemetry-javaagent.jar (JVM agent) |
| .NET | OpenTelemetry.AutoInstrumentation NuGet + env vars |
| Go | eBPF-based auto-instrumentation (experimental) |
| Ruby | opentelemetry-instrumentation-all gem |
1.3 Manual Span Creation Pattern
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
def process_order(order_id: str):
with tracer.start_as_current_span("process_order") as span:
span.set_attribute("order.id", order_id)
span.set_attribute("order.source", "api")
# Nested span for a sub-operation
with tracer.start_as_current_span("validate_inventory"):
check_inventory(order_id)
with tracer.start_as_current_span("charge_payment"):
charge(order_id)
span.set_status(trace.Status(trace.StatusCode.OK))
1.4 Context Propagation (W3C TraceContext)
All OpenTelemetry SDKs propagate trace context via W3C TraceContext headers by default:
traceparent: 00-{trace-id}-{parent-span-id}-{trace-flags}
tracestate: vendor-specific=value
Multi-service propagation is automatic when:
- HTTP clients are instrumented (auto-injection of headers)
- Message queues use OTel propagators
- All services use the same OTel exporter endpoint
Custom propagation for non-HTTP transports:
from opentelemetry.propagate import inject, extract
# Inject trace context into carrier (dict, message headers, etc.)
carrier = {}
inject(carrier)
kafka_headers = carrier # pass to Kafka message
# Extract on consumer side
ctx = extract(kafka_headers)
with tracer.start_as_current_span("consume", context=ctx):
process_message()
2. Monitoring & Metrics
2.1 RED vs USE Methodology
RED (Rate, Errors, Duration) — for Services
| Metric | Signal | Prometheus Example |
|---|---|---|
| Rate | Requests per second | rate(http_requests_total[5m]) |
| Errors | Failed request rate | rate(http_requests_total{status=~"5.."}[5m]) |
| Duration | Latency distribution | histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) |
RED applies to: HTTP APIs, gRPC services, worker pools, any request-driven service.
USE (Utilization, Saturation, Errors) — for Resources
| Metric | Signal | Prometheus Example |
|---|---|---|
| Utilization | % resource used | node_cpu_seconds_total{mode="idle"} → 100 - rate(...) |
| Saturation | Queue depth / load | node_load1, node_memory_SwapFree_bytes |
| Errors | Hardware/OS errors | node_network_receive_errs_total |
USE applies to: CPUs, memory, disks, network interfaces, database connection pools.
2.2 Prometheus Metric Types and Usage
# Counter — only ever increases (request count, errors)
# Functions: rate(), increase(), irate()
http_requests_total{method="GET", status="200"} 1023847
# Gauge — can go up and down (memory, queue depth, temp)
# Functions: avg_over_time(), max_over_time(), delta()
process_resident_memory_bytes 1.342e+08
# Histogram — bucketed observations (latency, size)
# Functions: histogram_quantile(), histogram_avg()
http_request_duration_seconds_bucket{le="0.1"} 450
http_request_duration_seconds_bucket{le="0.5"} 890
http_request_duration_seconds_bucket{le="+Inf"} 1000
http_request_duration_seconds_sum 1234.5
http_request_duration_seconds_count 1000
# Summary — client-side quantile computation (less flexible than histograms)
# Prefer histograms in most cases.
2.3 Cardinality Management
Cardinality = number of unique label combinations. High cardinality kills Prometheus.
DO:
- Keep label values bounded (<100 unique values):
status_code,http_method,endpoint - Use
droprelabel configs for noisy labels - Pre-aggregate in the Collector:
batch+memory_limiterprocessors
DON'T:
- ❌ Put user IDs, session IDs, or request IDs as labels
- ❌ Use unbounded dynamic values (timestamps, IPs, full URLs)
- ❌ Let GraphQL query names explode cardinality
Relabel example to drop high-cardinality labels:
relabel_configs:
- source_labels: [__name__]
regex: 'http_request_duration_seconds_bucket'
action: drop
# Drop if url label is set (too many unique values)
- source_labels: [url]
regex: '.+'
action: labeldrop
2.4 Recording Rules (Pre-computation)
# rules/recording_rules.yml
groups:
- name: http_aggregates
interval: 30s
rules:
- record: job:http_requests_total:rate5m
expr: rate(http_requests_total[5m])
- record: job:http_request_errors:rate5m
expr: rate(http_requests_total{status=~"5.."}[5m])
- record: job:http_request_duration:p99
expr: histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
- name: slo_dashboard
interval: 30s
rules:
- record: slo:error_budget_remaining:ratio
expr: |
1 - (
sum(rate(http_requests_total{status=~"5.."}[30d]))
/
sum(rate(http_requests_total[30d]))
) / 0.01 # 99% SLO
2.5 Grafana Dashboard Design
Golden signals dashboard layout:
| Row | Panels | Type |
|---|---|---|
| 1 | Request Rate + Error Rate | Graph (timeseries) |
| 2 | Latency p50/p90/p99 | Graph (timeseries) |
| 3 | Error Budget Remaining | Stat (gauge) |
| 4 | Top-N Endpoints by Latency | Table |
| 5 | Resource USE (CPU/Mem/Disk) | Graph (timeseries) |
| 6 | SLO Compliance (by endpoint) | Bar gauge |
3. Structured Logging
3.1 JSON Structured Logging Patterns
{
"timestamp": "2026-06-18T01:30:00.123Z",
"level": "info",
"message": "Order processed successfully",
"service": "order-service",
"trace_id": "0af7651916cd43dd8448eb211c80319c",
"span_id": "b7ad6b7169203331",
"order_id": "ORD-12345",
"customer_id": "CUST-789",
"duration_ms": 234,
"http": {
"method": "POST",
"path": "/api/orders",
"status_code": 201
}
}
3.2 Log Level Guidelines
| Level | Meaning | When to Use |
|---|---|---|
| ERROR | Operation failed; needs human attention | Unhandled exceptions, payment failures, data loss |
| WARN | Something unexpected; recoverable | Retry exhaustion, degraded mode, deprecation |
| INFO | Key business events; normal operation | Order created, user registered, deployment |
| DEBUG | Detailed troubleshooting info | Request payload, SQL queries, cache hits/misses |
| TRACE | Extremely verbose; line-level detail | Function entry/exit, variable dumps |
3.3 Trace Correlation
Every log line MUST include trace_id and span_id when inside a traced span. This enables single-click log-to-trace correlation in Grafana/Datadog.
Auto-injection patterns:
# Python: opentelemetry-instrumentation-logging auto-injects trace context
import logging
from opentelemetry.instrumentation.logging import LoggingInstrumentor
LoggingInstrumentor().instrument(set_logging_format=True)
# Now all log lines include:
# [2026-06-18 01:30:00,123] [INFO] [trace_id=0af7... span_id=b7ad...] message
// Node.js: Winston transport with OTel context
import { trace } from '@opentelemetry/api';
import winston from 'winston';
const logger = winston.createLogger({
format: winston.format.combine(
winston.format((info) => {
const span = trace.getActiveSpan();
if (span) {
info.trace_id = span.spanContext().traceId;
info.span_id = span.spanContext().spanId;
}
return info;
})(),
winston.format.json()
),
});
3.4 Log Aggregation (Loki)
Loki + Promtail pipeline:
# promtail-config.yml — scrape Kubernetes container logs
scrape_configs:
- job_name: kubernetes-pods
kubernetes_sd_configs:
- role: pod
pipeline_stages:
- json:
expressions:
level: level
trace_id: trace_id
service: service
- labels:
level:
service:
- output:
source: message
LogQL queries:
# Errors with trace correlation
{service="order-service", level="error"} | json | line_format "{{.message}}"
# Errors in the last hour, grouped by endpoint
sum by (http_path) (count_over_time({service="api-gateway"} | json | level="error" [1h]))
3.5 PII Redaction
# OTel Collector redaction processor
processors:
redaction:
allow_all_keys: false
allowed_keys:
- trace_id
- span_id
- service
- level
- message
- duration_ms
blocked_values:
- '.*@.*' # Email addresses
- '\d{3}-\d{2}-\d{4}' # SSN patterns
- '\b\d{16}\b' # Credit card numbers
4. Distributed Tracing
4.1 Trace Context Propagation Architecture
Client API Gateway Order Service Payment Service
| | | |
|--- HTTP GET -------->| | |
| traceparent=... | | |
| |--- gRPC call ------->| |
| | traceparent=... | |
| | |--- Kafka msg -------->|
| | | traceparent=... |
| | | in message headers |
4.2 Sampling Strategies
| Strategy | Description | When to Use | Config |
|---|---|---|---|
| AlwaysOn | 100% of traces | Development, low-volume | sampler=always_on |
| AlwaysOff | 0% of traces | Testing, no telemetry needed | sampler=always_off |
| Probability | Fixed % of traces | Stable production (e.g., 10%) | OTEL_TRACES_SAMPLER=traceidratio OTEL_TRACES_SAMPLER_ARG=0.1 |
| Rate limiting | Max N traces/sec | High-throughput services | sampler=rate_limiting |
| Parent-based | Follow parent's decision | Downstream services (default) | sampler=parentbased_always_on |
| Tail-based | Decision after span completes | Keep all errors + slow traces | Collector-level (load-balancing exporter) |
Recommended production config:
# OTel Collector tail sampling — keep all errors + >1s latency
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: latency
type: latency
latency: {threshold_ms: 1000}
- name: probabilistic
type: probabilistic
probabilistic: {sampling_percentage: 10}
4.3 Span Attributes Best Practices
# DO: Use semantic conventions
span.set_attribute("http.method", "POST")
span.set_attribute("http.status_code", 201)
span.set_attribute("db.system", "postgresql")
span.set_attribute("db.operation", "INSERT")
# DO: Add business context
span.set_attribute("order.value", 99.95)
span.set_attribute("order.items_count", 3)
# DON'T: High-cardinality attributes
# ❌ span.set_attribute("user.email", email)
# ❌ span.set_attribute("request.id", uuid4())
# ✔ Use span events for unique identifiers:
span.add_event("order_created", {"order_id": "ORD-12345"})
4.4 Error Recording
from opentelemetry.trace import Status, StatusCode
try:
result = process_order(order_id)
span.set_status(Status(StatusCode.OK))
except Exception as e:
span.set_status(Status(StatusCode.ERROR, str(e)))
span.record_exception(e, attributes={"order_id": order_id})
raise
4.5 Service Maps
Service maps are auto-generated by OTel backends (Grafana Tempo, Jaeger, Datadog) when trace context is consistently propagated across all services. Key requirements:
- Every service MUST propagate trace context to downstream calls
- Every service MUST export spans to the same collector/backend
- Span names should follow semantic conventions for proper grouping
5. Semantic Conventions
5.1 Span Naming
<resource>.<operation> — e.g., "HTTP GET", "gRPC OrderService/PlaceOrder"
<db.operation> <db.name> — e.g., "SELECT users", "INSERT orders"
<messaging.operation> <messaging.destination> — e.g., "process orders.new"
5.2 HTTP Semantic Conventions
| Attribute | Type | Example | Required |
|---|---|---|---|
http.method | string | GET, POST | Yes |
http.status_code | int | 200, 404 | Yes (if available) |
http.route | string | /users/:id | Recommended |
http.url | string | https://api.example.com/users/123 | Yes (client) |
http.target | string | /users/123?page=1 | Yes (server) |
http.request_content_length | int | 1024 | Optional |
http.response_content_length | int | 2048 | Optional |
network.protocol.version | string | 1.1, 2 | Recommended |
5.3 Database Semantic Conventions
| Attribute | Type | Example |
|---|---|---|
db.system | string | postgresql, mongodb, redis |
db.operation | string | SELECT, INSERT, find |
db.name | string | users_db |
db.statement | string | SELECT * FROM users WHERE id = ? |
db.mongodb.collection | string | orders |
db.redis.database_index | int | 0 |
5.4 Messaging Conventions
| Attribute | Type | Example |
|---|---|---|
messaging.system | string | kafka, rabbitmq, sqs |
messaging.operation | string | process, receive, publish |
messaging.destination | string | orders.new |
messaging.kafka.consumer_group | string | order-processor |
messaging.kafka.partition | int | 3 |
messaging.message.id | string | msg-12345 |
6. SLI / SLO / SLA
6.1 Definitions
| Term | Definition | Example | Owner |
|---|---|---|---|
| SLI | Service Level Indicator — the metric | "Ratio of successful requests to total requests" | Engineering |
| SLO | Service Level Objective — the target | "99.9% of requests succeed over 30 days" | Product + Eng |
| SLA | Service Level Agreement — the contract | "99.5% uptime or 10% credit" | Legal + Business |
6.2 SLI Types
Availability SLI
Good: HTTP 200-499 (non-5xx)
Bad: HTTP 5xx, timeouts, connection refused
SLI = good_requests / total_requests
Latency SLI
Good: requests completing within threshold (e.g., <300ms)
Bad: requests exceeding threshold
SLI = fast_requests / total_requests
Freshness SLI
Good: data processed within freshness window (e.g., <5min stale)
Bad: data older than freshness window
SLI = fresh_data_points / total_data_points
Coverage SLI
Good: data that passed validation/filtering
Bad: data dropped/ignored
SLI = processed_data / total_ingested_data
6.3 Error Budget
Error Budget = 1 - SLO_target
For 99.9% SLO over 30 days:
Total minutes: 43,200
Allowed downtime: 43.2 minutes/month
Error budget: 0.1%
Burn rate = actual_error_rate / budgeted_error_rate
A burn rate of 1: consuming budget at exactly the SLO pace
A burn rate of 10: consuming budget 10x faster than allowed
6.4 Multi-Window Burn Rate Alerts
# Prometheus alerting rules for burn rate alerts
groups:
- name: slo_burn_rate
rules:
# Fast burn: significant event, page on-call
- alert: SLOErrorBudgetBurnCritical
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
) > (0.01 * 14.4) # 1% budget, 14.4x burn rate = 1h
for: 2m
labels:
severity: critical
annotations:
summary: "Error budget burning 14.4x: 1h to exhaustion"
runbook: "https://runbooks.example.com/slo-burn-critical.md"
# Slow burn: warning, create ticket
- alert: SLOErrorBudgetBurnWarning
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[6h]))
/
sum(rate(http_requests_total[6h]))
) > (0.01 * 3) # 1% budget, 3x burn rate = 6h
for: 5m
labels:
severity: warning
annotations:
summary: "Error budget burning 3x: 6h window exceeded"
runbook: "https://runbooks.example.com/slo-burn-warning.md"
6.5 SLO Dashboard JSON Pattern
See scripts/generate-slo-dashboard.sh for automated dashboard generation from SLI definitions.
7. Alerting
7.1 Alert Design Principles
- Alert on symptoms, not causes — Alert on "user-facing error rate > 0.1%" not "CPU > 80%"
- Every alert must have a runbook — No runbook = no alert
- Eliminate toil alerts — Automate the response or remove the alert
- Page on SLO breaches only — Everything else can be a ticket/chat notification
- Test alerts regularly — Chaos engineering, fire drills, GameDays
7.2 Severity Classification
| Severity | Label | Response | Example |
|---|---|---|---|
| SEV0 | Critical | Page on-call immediately, 5min ack | Complete outage, data loss, SLO budget exhausted in <1h |
| SEV1 | High | Page on-call, 30min ack | Major feature broken, >50% error rate, budget burning at 10x |
| SEV2 | Medium | Create ticket, SLA 4h response | Single endpoint degraded, slow burn rate detected |
| SEV3 | Low | Create ticket, SLA 24h response | Non-critical component issue, capacity warning |
| SEV4 | Info | No action needed | Deprecation notice, planned maintenance |
7.3 Alert Routing (Alertmanager)
# alertmanager.yml
route:
receiver: 'default'
routes:
- match:
severity: critical
receiver: 'on-call-pager'
repeat_interval: 5m
group_wait: 10s
- match:
severity: warning
receiver: 'engineering-slack'
repeat_interval: 1h
- match_re:
service: '(order|payment).*'
receiver: 'payments-team'
receivers:
- name: 'on-call-pager'
pagerduty_configs:
- routing_key: 'your-pagerduty-key'
severity: critical
- name: 'engineering-slack'
slack_configs:
- channel: '#alerts-eng'
title: '{{ .GroupLabels.alertname }}'
text: '{{ .CommonAnnotations.summary }}'
- name: 'payments-team'
webhook_configs:
- url: 'https://hooks.slack.com/services/T...'
7.4 On-Call Rotation
# PagerDuty / Opsgenie escalation policy pattern
# Level 1: Primary on-call (5 min ack)
# Level 2: Secondary on-call (10 min ack, auto-escalate if L1 doesn't ack)
# Level 3: Engineering manager (30 min ack)
# Key practices:
# - Rotations should be at least 1 week (not daily)
# - Never have a single point of failure in the rotation
# - Shadow rotations for new on-call engineers
# - Post-on-call writeup within 24h of rotation end
7.5 Alert Fatigue Prevention
- Remove flapping alerts immediately — If it fires and resolves 5x in an hour, it's broken
- Aggregate during incidents — Group related alerts, don't page for every instance
- Tune thresholds quarterly — Review false positive rates
- "Business hours only" for SEV2 and below — Don't wake people up for non-urgent issues
- Inhibit alerts — Don't page for payment-service down if the network is down
# Alertmanager inhibition rule
inhibit_rules:
- source_match:
alertname: 'NetworkPartition' # Don't alert on...
target_match_re:
alertname: '.*Down' # ...anything-down if network is partitioned
equal: ['datacenter']
8. Incident Response
8.1 Incident Severity Levels
| Level | Description | Response Time | Communication Cadence |
|---|---|---|---|
| SEV0 | Full outage, data loss, security breach | Immediate | Every 30 min |
| SEV1 | Major functionality broken, high error rate | 5 min | Every 1 hour |
| SEV2 | Partial degradation, single feature affected | 30 min | Every 4 hours |
| SEV3 | Minor issue, no user impact | 4 hours | Status page update |
8.2 Incident Commander (IC) Role
The IC is responsible for coordination, NOT necessarily fixing the problem.
IC responsibilities:
- Declare the incident and severity
- Set up the incident channel (Slack/Zoom)
- Assign roles: Ops Lead, Comms Lead, Scribe
- Maintain the incident timeline
- Decide when to escalate
- Declare incident resolved
- Schedule and lead the postmortem
8.3 Communication Templates
Incident Declaration (Slack)
🚨 INCIDENT DECLARED: {title}
Severity: {SEV0/SEV1/SEV2}
IC: {name}
Ops Lead: {name}
Incident Channel: #{channel}
Zoom: {link}
Summary: {one-line description of what's happening}
Customer Impact: {who is affected and how}
Start Time: {ISO timestamp}
Status Update (Every 30-60 min)
📊 INCIDENT UPDATE #{N}: {title}
Time elapsed: {duration}
Status: {investigating/mitigating/resolved}
Current understanding:
- {bullet point findings}
Actions taken:
- {bullet point actions}
Next steps:
- {bullet point next actions}
ETA to resolution: {estimate}
Incident Resolution
✅ INCIDENT RESOLVED: {title}
Duration: {start_time} to {end_time} ({total_duration})
Severity: {SEV0/SEV1/SEV2}
Root Cause: {brief description}
Fix: {what was done to resolve}
Customer Impact: {final impact summary}
Postmortem: scheduled for {date} — {link}
Ticket: {ticket link}
8.4 Timeline Reconstruction Template
## Incident Timeline: {title}
| Time (UTC) | Event | Source | Actor |
|------------|-------|--------|-------|
| 14:00 | Deploy v2.4.1 started | Deployment tool | @engineer |
| 14:03 | Latency spike detected (>500ms) | Grafana alert | System |
| 14:05 | Alert fired: SLOErrorBudgetBurnCritical | Alertmanager | System |
| 14:07 | IC declared SEV1 | Slack | @ic-name |
| 14:12 | Identified deploy as trigger | Ops investigation | @ops-lead |
| 14:15 | Rollback initiated | CI/CD | @ops-lead |
| 14:18 | Metrics recovering | Grafana | System |
| 14:22 | Service fully recovered | Grafana | System |
| 14:30 | Incident resolved | Slack | @ic-name |
8.5 Postmortem Structure
# Postmortem: {incident title}
**Date:** YYYY-MM-DD
**Authors:** {names}
**Severity:** {SEV0/SEV1/SEV2}
**Duration:** {start → end, total duration}
## Summary
{2-3 sentence summary of what happened and impact}
## Customer Impact
- Who was affected and for how long
- What functionality was degraded/unavailable
- Error budget consumed: X% of monthly budget
## Timeline
{Same format as Section 8.4 — copy from incident channel}
## Root Cause Analysis
### Direct Cause
{The technical thing that broke}
### Contributing Factors
- {Why the direct cause was possible}
- {What allowed it to propagate}
- {What delayed detection}
## Detection
- How was it detected? (Alert, user report, social media)
- How long from start to detection? (TTD)
- How long from detection to resolution? (TTR)
- Could detection have been faster? How?
## Resolution
- What action resolved the incident?
- Was any data lost or corrupted?
## Action Items
| Priority | Action | Owner | Due |
|----------|--------|-------|-----|
| P0 | {critical fix to prevent recurrence} | @owner | YYYY-MM-DD |
| P1 | {improvement} | @owner | YYYY-MM-DD |
| P2 | {nice-to-have} | @owner | YYYY-MM-DD |
## Lessons Learned
- What went well
- What went poorly
- Where we got lucky (near-misses)
9. Observability as Code
9.1 Terraform: Grafana Dashboards + Alerts
# grafana-dashboard.tf
resource "grafana_dashboard" "service_overview" {
folder = grafana_folder.services.id
config_json = file("${path.module}/dashboards/service-overview.json")
}
resource "grafana_alert_rule" "error_rate" {
name = "High Error Rate - Order Service"
folder_uid = grafana_folder.alerts.uid
rule_group = "service-alerts"
for = "5m"
condition = "C"
no_data_state = "NoData"
exec_err_state = "Error"
# Query: error rate > 1%
queries {
ref_id = "A"
datasource_uid = "prometheus"
expr = "sum(rate(http_requests_total{service=\"order\",status=~\"5..\"}[5m])) / sum(rate(http_requests_total{service=\"order\"}[5m])) > 0.01"
}
annotations = {
runbook_url = "https://runbooks.example.com/order-service-errors.md"
}
labels = {
severity = "critical"
}
}
9.2 GitOps Workflow for Monitoring Config
monitoring-config/
├── dashboards/
│ ├── service-overview.json
│ ├── slo-compliance.json
│ └── infrastructure-overview.json
├── alerts/
│ ├── slo-burn-rate.yml
│ ├── infrastructure.yml
│ └── application.yml
├── rules/
│ ├── recording-rules.yml
│ └── silencers.yml
├── terraform/
│ ├── main.tf
│ └── variables.tf
└── .github/workflows/
└── deploy-monitoring.yml
GitOps workflow:
- PR to change dashboard/alert → code review
- Merge to main → CI runs
promtool check rules+ dashboard JSON validation - CI applies via Terraform to Grafana/Prometheus
- Drift detection cron job reconciles every hour
10. Cost Optimization
10.1 Cardinality Management Checklist
- Audit metric label cardinality monthly
- Set
max_cardinalitylimits on high-risk dimensions - Use
droprelabel configs for unused labels - Pre-aggregate with recording rules (reduce raw data retention)
- Monitor
prometheus_tsdb_head_seriesfor growth trends
10.2 Sampling Cost Calculator
Annual trace storage cost = traces_per_second * avg_spans_per_trace
* avg_span_size_bytes * 86400 * 365 * sampling_rate * $per_GB
Example (head sampling at 10%):
1000 req/s * 10 spans * 1KB * 86400 * 365 * 0.10 * $0.50/GB
= 1000 * 10 * 1024 * 86400 * 365 * 0.10 * 0.0000000005
≈ $16,181/year
Example (tail sampling at 1% with error/slow retention):
Same base but keep 1% normal + 100% errors + 100% slow (>1s)
If 5% errors and 2% slow, total retained ≈ 8%
≈ $12,945/year — savings of 20%
10.3 Retention Policies
| Data Type | Hot Storage | Warm Storage | Cold Storage | Rationale |
|---|---|---|---|---|
| Metrics (raw) | 7 days | 30 days | — | High volume, fast query is key |
| Metrics (aggregated) | 30 days | 90 days | 1 year | For capacity planning, trends |
| Traces | 3 days | 14 days | — | Debugging window; sample for long-term |
| Logs | 7 days | 30 days | 90 days | Compliance often requires longer |
11. Quick-Start Checklists
Production Readiness Checklist
- Auto-instrumentation enabled for all services
- Manual spans for business-critical operations
- Trace context propagated across all service boundaries
- RED metrics dashboards for all user-facing services
- USE metrics dashboards for all infrastructure
- Structured JSON logging with trace_id in every log line
- SLOs defined and SLO dashboards published
- Burn rate alerts configured (fast + slow burn)
- Alert routing tested end-to-end
- Runbooks linked in every alert annotation
- Incident response playbook documented
- On-call rotation configured and tested
- Cardinality audit completed
- Sampling strategy reviewed and documented
- Dashboard JSON validated in CI
- Alert rules syntax-checked in CI
Debugging with Observability (Troubleshooting Flow)
- Start with the alert → Which SLO is burning? Which service?
- Check the SLO dashboard → Isolate the failing endpoint or dependency
- Look at traces → Find a representative failing trace, follow the waterfall
- Correlate with logs → Click from trace span to logs (via trace_id)
- Check recent deploys → Overlay deployment markers on dashboards
- Check dependent service SLOs → Is the failure upstream?
- Post-incident → Update runbook, file action items from postmortem
References
references/otel-instrumentation-guide.md— Multi-language instrumentation deep-divereferences/sli-slo-cookbook.md— SLI patterns, error budgets, burn rate configsreferences/incident-response.md— Full incident management playbook- OpenTelemetry Specification
- Prometheus Alerting Rules
- Google SRE Book — SLO Chapter