agentsclimarketplace

Set up drift alerts

Skill ContextJet-ai/awesome-llm-observability/skills/set-up-drift-alerts

Use this to catch an LLM app silently getting worse in production - quality dropping, cost creeping up, inputs shifting away from what you tested. Trigger on "monitor my LLM in production", "alert me when quality drops", "detect drift", "my app got worse and I didn't notice", "set up monitoring/alerting for my AI app". Alert on the signals that actually move, not vanity metrics.From its SKILL.md

Install
npx -y skills add ContextJet-ai/awesome-llm-observability --skill set-up-drift-alerts

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

What its file declares

Copied from the file, not written here

The file declares its own license as CC0-1.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

3.5 KB, 761 tokens by cl100k_base, as published. Nobody here has run it

Set up drift alerts

An LLM app that worked at launch degrades quietly: a model provider updates weights, users start asking different questions, a prompt change trades quality for cost. Drift alerting turns "we found out from an angry customer" into "we got paged when the metric moved."

The four drifts worth alerting on

  1. Quality drift - your online eval scores (faithfulness, relevance, hallucination rate) trend down. The one that matters most and the one people forget to watch. Requires online evals (see add-llm-evals).
  2. Cost drift - tokens/request or spend/day creeps up (context bloat, a retry bug, a prompt that grew). Cheap to track, common, expensive to ignore.
  3. Latency / error drift - p95 latency or error rate rises (provider slowdown, rate limits, a slow retrieval).
  4. Input drift - production inputs move away from your eval set (new topics, new languages, longer inputs). Your evals stop being representative, so quality can drop without your quality metric catching it. Detect via input length/embedding-distribution shift.

How to set it up

  1. Emit the signals as span attributes / metrics (see instrument-llm-observability): per-request tokens, cost, latency, error, and online eval score.
  2. Baseline them over a stable window (e.g. last 7-14 days) to know "normal."
  3. Alert on deviation, not absolute values - page when a metric moves meaningfully vs its own baseline (e.g. quality score down >X%, cost/request up >Y%, error rate over threshold). Absolute thresholds age badly; relative-to-baseline survives growth.
  4. Route alerts where the team actually looks (Slack/PagerDuty), with the trace/dashboard link attached so triage is one click.
  5. Sample, don't score everything - online eval on a % of traffic is enough to see a trend and keeps cost sane.

Most observability platforms (Langfuse, Phoenix, Opik, Datadog LLM Obs) have built-in dashboards + alerting for these - wire the signals in and set the thresholds rather than building from scratch.

Verify

  • Deliberately regress a prompt in staging and confirm the quality alert fires.
  • Confirm cost/latency/error alerts fire against a baseline, not a hard-coded number.
  • Confirm the alert lands where the team sees it, with a link to the offending traces.

Anti-patterns

  • Alerting on cost/latency but never on quality (the app silently gets dumber while staying cheap and fast).
  • Absolute thresholds that either page constantly or never (tie alerts to a rolling baseline).
  • Scoring 100% of traffic with LLM-as-judge (expensive; sample instead).
  • Alerts nobody sees, or with no link to the trace (they get muted, then ignored).

Grounding

Drift/monitoring for ML has a long lineage (data & concept drift); the LLM-specific additions are online LLM-as-a-judge quality scoring (Zheng et al. 2023, arXiv:2306.05685) and token/cost/latency telemetry via the OpenTelemetry GenAI semantic conventions. Open-source drift/eval monitoring: Evidently.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most monitoring observability skills give in 761 tokens

Counted across 530 of the 532 authors here whose files we hold, read 2026-09-06

  • Use structured JSON loggingin 40 of 530, across 36 files
  • Link every alert to a runbookin 29 of 530, across 27 files
  • Attach correlation IDs to every log linein 19 of 530, across 16 files
  • Alert on symptoms rather than causesin 19 of 530, across 17 files
  • Use OpenTelemetry for distributed tracingin 15 of 530, across 14 files
  • Alert on symptoms users feelin 15 of 530, across 13 files
  • Implement health check endpointsin 14 of 530, across 10 files
  • Inspect existing dashboards firstin 12 of 530, across 4 files
  • Build the minimum useful boardin 12 of 530, across 4 files
  • Start from operator questionsin 12 of 530, across 4 files
  • Propagate trace context across boundariesin 11 of 530, across 10 files
  • Include trace id in all log entriesin 10 of 530, across 9 files

Said here and by no other author read

  • Emit the signals as span attributes
  • Baseline them over a stable window
  • Alert on deviation from the baseline
  • Route alerts to team communication channels
  • Sample traffic for online evaluation
  • Regress a prompt in staging to verify

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.