Observability stack
Skill sairam0424/MindForge/.mindforge/skills/observability-stack
MindForge: The Enterprise Agentic Framework for Claude Code & Antigravity. High-performance autonomous execution, wave-parallelism, and multi-tier governance for production-grade AI engineering.
npx -y skills add sairam0424/MindForge --skill observability-stackAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.6 KB, ~1.7k tokens by cl100k_base, as published. Nobody here has run it
Observability Stack
When this skill activates
This skill activates when designing, implementing, or reviewing observability infrastructure for production systems. It covers the three pillars (logs, traces, metrics), alerting strategy, dashboard design, and correlation techniques. Use this skill whenever you need to ensure a system can be understood, debugged, and monitored in production.
Mandatory actions when this skill is active
Before
- Identify observability goals — What questions must you answer? "Why is this request slow?" (traces), "What's the error rate?" (metrics), "What happened at 3:42am?" (logs).
- Assess current state — What instrumentation exists? What gaps cause blind spots? Never add observability without understanding what's already there.
- Define SLIs/SLOs — Service Level Indicators (what to measure) and Service Level Objectives (acceptable thresholds) must exist before alerting makes sense.
- Choose the stack — Select instrumentation (OpenTelemetry), storage (Prometheus/Loki/Tempo, Datadog, Grafana Cloud), and visualization (Grafana, Datadog) based on team expertise and budget.
During
Pillar 1: Structured Logging
Format — Always structured JSON with fields: timestamp, level, message, service, trace_id, span_id, correlation_id, error_code, duration_ms, metadata.
Severity levels:
DEBUG— Development-only detail. Never enable in production by default.INFO— Normal operations: request received, task completed, state transitions.WARN— Degraded but functional: retry succeeded, fallback activated, approaching limit.ERROR— Operation failed: request could not be fulfilled, exception caught.FATAL— System cannot continue: startup failure, unrecoverable state, crash.
Correlation IDs:
- Generate at the edge (API gateway). Propagate via
X-Correlation-IDheader through all services. - Include in every log line. Distinct from trace_id but can be linked.
Logging rules:
- Log at service boundaries (incoming requests, outgoing calls).
- Log state transitions (order created, payment confirmed, email sent).
- Log errors with full context (what was attempted, what failed, what the input was).
- NEVER log secrets, tokens, passwords, or full credit card numbers.
- NEVER log PII without a documented legal basis and redaction strategy.
Pillar 2: Distributed Tracing
Span design:
- Create spans for: HTTP requests (in/out), database queries, cache operations, external API calls, queue publish/consume, significant business logic.
- DO NOT span: trivial operations, tight loops, in-memory calculations.
- Name spans with:
verb.nounformat:process.payment,query.users,send.email.
Context propagation:
- Use W3C Trace Context standard (
traceparentheader) for cross-service propagation. - Ensure all HTTP clients and message queues propagate context automatically.
- Test propagation: a trace should span from the edge to the deepest downstream call.
Sampling strategy:
- 100% sampling in development and staging.
- Production: tail-based sampling (keep all errors + slow requests + 1-10% of normal traffic).
- Never sample at 100% in high-traffic production. Storage costs scale linearly.
- Always keep traces for errors regardless of sampling rate.
Pillar 3: Metrics
RED Method (for services/endpoints):
- Rate — Requests per second. Indicates load.
- Errors — Failed requests per second (and error rate as %). Indicates reliability.
- Duration — Latency distribution (p50, p95, p99). Indicates performance.
USE Method (for resources: CPU, memory, disk, network):
- Utilization — Percentage of resource capacity in use.
- Saturation — Work queued waiting for the resource.
- Errors — Resource-level errors (disk I/O errors, network drops).
Metric types:
- Counter — Monotonically increasing value. Use for: total requests, errors, bytes transferred.
- Gauge — Point-in-time value that can go up/down. Use for: active connections, queue depth, temperature.
- Histogram — Distribution of values. Use for: request latency, payload sizes.
Naming convention: <service>_<component>_<metric>_<unit> (e.g., payment_api_request_duration_seconds).
Cardinality discipline:
- Never use unbounded values as label/tag keys (user IDs, request IDs, timestamps).
- Maximum 10-20 unique values per label. High cardinality kills metric storage.
- Safe labels: HTTP method, status code, endpoint path (grouped), environment, region.
Alerting Design
Principles:
- Alert on symptoms (user impact), not causes (CPU spike). Users don't care about CPU; they care about latency.
- Every alert must be actionable. If no one can do anything about it at 3am, it should not page.
- Use SLO-based alerting: alert when the error budget burn rate exceeds safe levels.
Threshold design:
- Static thresholds: set at baseline + 2 standard deviations (for stable metrics).
- Anomaly detection: use ML-based for metrics with seasonal patterns (traffic).
- Burn rate: alert when 1-hour burn rate > 14x (consumes 24hr budget in 1hr).
Alert severity:
- Critical (page) — User-facing impact right now. Revenue loss. Data loss.
- Warning (ticket) — Degraded performance. Will become critical if unaddressed. Handle next business day.
- Info (log) — Notable but not impactful. For awareness only.
Anti-patterns:
- Alert fatigue: >5 alerts/day/team = too many. Tune or eliminate.
- Flapping alerts: add hysteresis (alert fires at threshold, clears at threshold - 10%).
- Duplicate alerts: one incident should fire one alert, not one per symptom.
Dashboard Design
Layout: Row 1: Traffic, Error rate, Latency (p50/p95/p99). Row 2: Saturation (CPU, memory, connections), Active alerts. Row 3: Deployment markers, SLO burn-down.
Principles: Top = current health (red/yellow/green). Middle = golden signal time-series. Bottom = drill-down investigation panels. NO vanity metrics. Consistent time ranges, default 1 hour.
After
- Verify correlation — Generate a test request and verify it can be traced from edge to database across all services via correlation ID and trace ID.
- Validate alerting — Trigger each alert condition deliberately. Verify it fires, routes correctly, and auto-resolves.
- Load test observability — Ensure the observability stack itself handles production load without becoming a bottleneck.
- Document runbooks — Every alert needs a runbook: what it means, how to diagnose, how to mitigate.
Self-check before task completion
- All logs are structured JSON with correlation IDs, timestamps, and severity levels
- Distributed traces span from edge to deepest downstream call with proper context propagation
- RED metrics (Rate, Errors, Duration) are collected for every service endpoint
- USE metrics (Utilization, Saturation, Errors) are collected for critical resources
- Metric label cardinality is bounded (no unbounded labels like user IDs)
- Alerts are symptom-based, actionable, and severity-appropriate
- Dashboard shows golden signals with drill-down capability
- No secrets or unbounded PII appear in logs
- Sampling strategy is defined for production tracing
- Runbooks exist for every critical alert
Gives 2 of the 12 instructions most monitoring observability skills give in ~1.7k tokens
Counted across 481 of the 483 authors here whose files we hold, read 2026-08-06
- link every alert to a runbookin 43 of 481, across 35 files
- use structured json logginghere, and in 36 of 481, across 31 files
- alert on user-facing symptomsin 20 of 481, across 15 files
- emit structured JSON logs with stable event namesin 18 of 481, across 13 files
- propagate trace context across boundariesin 16 of 481
- use histograms for latency trackingin 14 of 481, across 9 files
- use OpenTelemetry for distributed tracingin 13 of 481, across 8 files
- include a correlation ID on every log linein 13 of 481, across 8 files
- Define service level objectiveshere, and in 10 of 481, across 7 files
- Call useAzureMonitor before importing other modulesin 9 of 481, across 2 files
- stop and ask for clarification if inputs are missingin 9 of 481, across 2 files
- define on-call questions before adding telemetryin 9 of 481, across 4 files
Said here and by no other author read
- choose observability tools based on team expertise
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.