Grafana
Query and manage Grafana dashboards, alert rules, and data sources via HTTP API. Use when viewing dashboards, troubleshooting alerts, checking service metrics, finding data sources, or when Grafana, monitoring, alerts, dashboards, or observability is mentioned.From its SKILL.md
npx -y skills add addxai/enterprise-harness-engineering --skill grafanaAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- reads credentialsReads from 1 credential source: `GRAFANA_TOKEN`.
SKILL.md
3.6 KB, 914 tokens by cl100k_base, as published. Nobody here has run it
Grafana
Query and manage Grafana monitoring dashboards, alert rules, and data sources via HTTP API.
Setup
Configure your Grafana instance:
| Variable | Description | Required |
|---|---|---|
GRAFANA_URL | Your Grafana server URL (e.g., https://grafana.example.com) | Yes |
GRAFANA_TOKEN | API Key (Settings → API Keys, viewer or editor role) | Yes |
Authentication: curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" "$GRAFANA_URL/api/..."
Customize the following for your organization:
- Folder structure: How dashboards are organized (by team, region, service, etc.)
- Tag conventions: What tags are used for filtering (region, environment, service type)
- Alert naming: Your alert naming pattern (e.g.,
{Service} {Metric} {Condition} [{env}]) - Key dashboards: Which dashboards to check after deployments
Rules
Dashboard Operations
- Search first, then detail: Use search API with tags/query to narrow scope, then fetch by uid
- Never delete production dashboards — archive by moving to an Archive folder
- Write operations require user confirmation before execution (create/modify dashboards, alert rules)
Common Workflows
Find a Dashboard:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/search?query=<keyword>&tag=<tag>&type=dash-db" \
| jq '.[] | {uid, title, folderTitle}'
Get Dashboard Details:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/dashboards/uid/<uid>" \
| jq '.dashboard.panels[] | {title, type}'
Check Active Alerts:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/alerts?state=alerting"
List Data Sources:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/datasources" \
| jq '.[] | {id, name, type, url}'
Search by Folder:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/search?folderIds=<folder_id>&type=dash-db"
Post-Deployment Checks
After each deployment, check key dashboards for:
- Error rate changes (before vs after deployment)
- Latency P95/P99 trends
- Consumer lag (if using message queues)
- Connection counts (TCP/WebSocket)
- Log level distribution (error/warn spikes)
Alert Troubleshooting
GET /api/alerts?state=alerting— list all firing alerts- Identify the dashboard and panel from the alert
- Read the PromQL expression from the panel
- Query Prometheus directly to understand the data
- Check recent deployments or config changes as potential cause
Examples
Bad
# Delete a production dashboard (never delete, archive instead)
curl -X DELETE "$GRAFANA_URL/api/dashboards/uid/abc123"
# Fetch all dashboards without filtering (use search first)
for uid in $(curl ... /api/search | jq -r '.[].uid'); do
curl ... /api/dashboards/uid/$uid
done
Good
# Search dashboards by tag
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/search?tag=production&type=dash-db" \
| jq '.[] | {uid, title, folderTitle}'
# Check active alerts
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/alerts?state=alerting"
# Get alert details for troubleshooting
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_URL/api/alerts/<alert_id>"
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most monitoring observability skills give in 914 tokens
Counted across 530 of the 532 authors here whose files we hold, read 2026-09-06
- Use structured JSON loggingin 40 of 530, across 36 files
- Link every alert to a runbookin 29 of 530, across 27 files
- Attach correlation IDs to every log linein 19 of 530, across 16 files
- Alert on symptoms rather than causesin 19 of 530, across 17 files
- Use OpenTelemetry for distributed tracingin 15 of 530, across 14 files
- Alert on symptoms users feelin 15 of 530, across 13 files
- Implement health check endpointsin 14 of 530, across 10 files
- Inspect existing dashboards firstin 12 of 530, across 4 files
- Build the minimum useful boardin 12 of 530, across 4 files
- Start from operator questionsin 12 of 530, across 4 files
- Propagate trace context across boundariesin 11 of 530, across 10 files
- Include trace id in all log entriesin 10 of 530, across 9 files
Said here and by no other author read
- Search first then detail
- Get user confirmation before write operations
- Archive dashboards instead of deleting
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.