Observability slo
Skill charlieviettq/awesome-agent-skill/.cursor/skills/reliability-ops/observability-slo
Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).
npx -y skills add charlieviettq/awesome-agent-skill --skill observability-sloAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 22 stars22 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Lightweight observability and SLO practice—SLIs, SLO targets, error budgets, logs, metrics, traces, and alerting for apps and data pipelines. Use when defining monitoring, on-call alerts, or reliability targets. Triggers: "SLO", "SLI", "observability", "alerting", "error budget", "monitoring".
SKILL.md
1.7 KB, as published. Nobody here has run it
Observability and SLO (lite)
Three pillars (minimum viable)
| Pillar | Start here |
|---|---|
| Logs | Structured JSON; request/job ID; severity |
| Metrics | RED/USE for services; duration, errors, throughput |
| Traces | One trace per request/job across critical hops |
SLI -> SLO flow
- Pick SLI — measurable user- or business-visible signal.
- Set SLO — target over window (e.g. 99.9% availability / 30d).
- Error budget — 100% - SLO; spend triggers policy (slow features, freeze risky deploys).
- Alert on budget burn — fast burn (pages) vs slow burn (ticket).
Example SLIs
| Service type | SLI examples |
|---|---|
| API | Success rate, p95 latency |
| Batch job | Completion within SLA window |
| Model scoring | Scoring latency, feature freshness |
Alert rules
- Page humans only for SLO-threatening or user-visible outages.
- Warning for degradation trending toward budget burn.
- Every alert links to a runbook (symptom, checks, mitigation, escalation).
Runbook skeleton
## Symptom
## Impact
## Checks (ordered)
## Mitigation
## Escalation
Pipeline note
For data/ML jobs: monitor row counts, null spikes, partition lag, and scoring drift alongside infra metrics.
Anti-patterns
- Alert on every log error without SLO linkage.
- SLOs without measurement (aspirational only).
- Dashboards nobody owns.