agentsclimarketplace

Grafana

Skill addxai/enterprise-harness-engineering/skills/grafana

Query and manage Grafana dashboards, alert rules, and data sources via HTTP API. Use when viewing dashboards, troubleshooting alerts, checking service metrics, finding data sources, or when Grafana, monitoring, alerts, dashboards, or observability is mentioned.From its SKILL.md

Install
npx -y skills add addxai/enterprise-harness-engineering --skill grafana

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • reads credentialsReads from 1 credential source: `GRAFANA_TOKEN`.

SKILL.md

3.6 KB, 914 tokens by cl100k_base, as published. Nobody here has run it

Grafana

Query and manage Grafana monitoring dashboards, alert rules, and data sources via HTTP API.

Setup

Configure your Grafana instance:

VariableDescriptionRequired
GRAFANA_URLYour Grafana server URL (e.g., https://grafana.example.com)Yes
GRAFANA_TOKENAPI Key (Settings → API Keys, viewer or editor role)Yes

Authentication: curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" "$GRAFANA_URL/api/..."

Customize the following for your organization:

  • Folder structure: How dashboards are organized (by team, region, service, etc.)
  • Tag conventions: What tags are used for filtering (region, environment, service type)
  • Alert naming: Your alert naming pattern (e.g., {Service} {Metric} {Condition} [{env}])
  • Key dashboards: Which dashboards to check after deployments

Rules

Dashboard Operations

  • Search first, then detail: Use search API with tags/query to narrow scope, then fetch by uid
  • Never delete production dashboards — archive by moving to an Archive folder
  • Write operations require user confirmation before execution (create/modify dashboards, alert rules)

Common Workflows

Find a Dashboard:

curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/search?query=<keyword>&tag=<tag>&type=dash-db" \
  | jq '.[] | {uid, title, folderTitle}'

Get Dashboard Details:

curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/dashboards/uid/<uid>" \
  | jq '.dashboard.panels[] | {title, type}'

Check Active Alerts:

curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/alerts?state=alerting"

List Data Sources:

curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/datasources" \
  | jq '.[] | {id, name, type, url}'

Search by Folder:

curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/search?folderIds=<folder_id>&type=dash-db"

Post-Deployment Checks

After each deployment, check key dashboards for:

  • Error rate changes (before vs after deployment)
  • Latency P95/P99 trends
  • Consumer lag (if using message queues)
  • Connection counts (TCP/WebSocket)
  • Log level distribution (error/warn spikes)

Alert Troubleshooting

  1. GET /api/alerts?state=alerting — list all firing alerts
  2. Identify the dashboard and panel from the alert
  3. Read the PromQL expression from the panel
  4. Query Prometheus directly to understand the data
  5. Check recent deployments or config changes as potential cause

Examples

Bad

# Delete a production dashboard (never delete, archive instead)
curl -X DELETE "$GRAFANA_URL/api/dashboards/uid/abc123"

# Fetch all dashboards without filtering (use search first)
for uid in $(curl ... /api/search | jq -r '.[].uid'); do
  curl ... /api/dashboards/uid/$uid
done

Good

# Search dashboards by tag
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/search?tag=production&type=dash-db" \
  | jq '.[] | {uid, title, folderTitle}'

# Check active alerts
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/alerts?state=alerting"

# Get alert details for troubleshooting
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
  "$GRAFANA_URL/api/alerts/<alert_id>"

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most monitoring observability skills give in 914 tokens

Counted across 530 of the 532 authors here whose files we hold, read 2026-09-06

  • Use structured JSON loggingin 40 of 530, across 36 files
  • Link every alert to a runbookin 29 of 530, across 27 files
  • Attach correlation IDs to every log linein 19 of 530, across 16 files
  • Alert on symptoms rather than causesin 19 of 530, across 17 files
  • Use OpenTelemetry for distributed tracingin 15 of 530, across 14 files
  • Alert on symptoms users feelin 15 of 530, across 13 files
  • Implement health check endpointsin 14 of 530, across 10 files
  • Inspect existing dashboards firstin 12 of 530, across 4 files
  • Build the minimum useful boardin 12 of 530, across 4 files
  • Start from operator questionsin 12 of 530, across 4 files
  • Propagate trace context across boundariesin 11 of 530, across 10 files
  • Include trace id in all log entriesin 10 of 530, across 9 files

Said here and by no other author read

  • Search first then detail
  • Get user confirmation before write operations
  • Archive dashboards instead of deleting

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.