agentsclimarketplace

Prometheus alerting cardinality review

Skill Raishin/vanguard-frontier-agentic/skills/prometheus/prometheus-alerting-cardinality-review

Curated marketplace of AI skills, agents, and rules for cloud, zero-trust, and compliance-aware engineering - works with Claude Code, Codex, Cursor, Copilot, and more.

Install
npx -y skills add Raishin/vanguard-frontier-agentic --skill prometheus-alerting-cardinality-review

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 18 stars18 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use this skill when reviewing Prometheus or AlertManager configuration for cardinality, alerting correctness, scrape security, remote_write safety, or retention adequacy. Trigger when a user provides prometheus.yml, alertmanager.yml, recording rules YAML, alerting rules YAML, or asks whether their Prometheus setup is production-ready.

SKILL.md

3.0 KB, as published. Nobody here has run it

Prometheus Alerting and Cardinality Review

Purpose

This skill reviews Prometheus and AlertManager configuration for cardinality explosion risks, recording rule adequacy, alert expression correctness, routing tree safety, scrape configuration security, and retention posture. Cardinality explosion is the leading cause of Prometheus OOM crashes in production, and flapping alerts from missing for: durations erode on-call trust faster than any other alerting defect.

Lean operating rules

  • Flag any label dimension that is unbounded at the application level (e.g., user_id, request_id, session_id, url_path, pod_hash) — these cause cardinality explosion and must be moved off the label set or aggregated away.
  • Treat prometheus_tsdb_head_series exceeding 5 million as a cardinality warning threshold; note it if the user reports series counts or if the config makes it likely.
  • Treat any alert rule with for: 0m, for: 0s, or no for: field as HIGH — bare threshold alerts flap on every scrape jitter.
  • Treat honor_labels: true on any scrape target that is not a trusted federation endpoint as HIGH — it allows the scraped workload to override job and instance labels.
  • Treat any scrape config with a non-cluster HTTP scheme (http://external-host) as a potential SSRF candidate and flag it.
  • Recording rules are required for any PromQL expression used in dashboards or SLO burn-rate calculations; flag their absence as MEDIUM.
  • Multi-window multi-burn-rate (MWMB) alerting is the correct pattern for SLO breach detection; flag single-window SLO alerts as MEDIUM.
  • Flag remote_write configs where write_relabel_configs drop non-__ metric labels — data loss is silent.
  • Flag retention under 30 days with no remote_write or Thanos/Cortex integration as MEDIUM compliance risk.
  • Do not recommend disabling any existing alert or recording rule without stating the specific reason and risk trade-off.

References

Load these only when needed:

Response minimum

Return, at minimum:

  • Cardinality risk assessment (label audit findings)
  • Alert expression correctness findings (for: duration, absent misuse, MWMB posture)
  • AlertManager routing and inhibition findings
  • Scrape config security findings
  • Retention and remote_write findings
  • Severity-labelled finding list (critical / high / medium / low)
  • Safe next actions

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.