agentsclimarketplace

Regression watch

Skill hoangsonww/Claude-Code-Agent-Monitor/plugins/ccam-insights/skills/regression-watch

πŸš€ A real-time monitoring dashboard for Claude Code, built with SQLite3, Node.js, Express, React, Vite, TailwindCSS, and WebSockets. It tracks sessions, agent activity, tool usage, and subagent orchestration, providing live analytics, a Kanban status board, status notifications, a cute buddy, and an interactive web UI/MacOS/Windows native app.

Install
npx -y skills add hoangsonww/Claude-Code-Agent-Monitor --skill regression-watch

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Detect quality and efficiency regressions over time using Agent Monitor data β€” rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.

SKILL.md

4.0 KB, as published. Nobody here has run it

Regression Watch

Detect whether Claude Code sessions are getting worse over time across quality and efficiency metrics, using Agent Monitor data.

Input

The user provides: $ARGUMENTS

This may be:

  • empty or "all" β€” check every regression metric (default)
  • "errors" β€” error-rate regression only
  • "cache" β€” cache hit-rate regression only
  • "compaction" β€” compaction-frequency regression only
  • "cost" β€” cost-per-session regression only
  • A window like "last 30d" or "30 vs 90" β€” set the recent vs baseline window sizes

Data Sources

EndpointReturns
GET /api/analyticsdaily_events (365d), daily_sessions (365d), event_types, tokens (total_input, total_output, total_cache_read, total_cache_write β€” baselines pre-summed), avg_events_per_session
GET /api/events?session_id=XEvent stream incl. APIError, Compaction, PreToolUse/PostToolUse β€” used to localize regressions to specific sessions
GET /api/pricing/cost{ total_cost, breakdown[...] } β€” total cost to derive cost-per-session
GET /api/pricing/cost/{sessionId}Per-session cost β€” used to compare recent vs baseline session cost
GET /api/workflows/{sessionId}compaction (impact), errorPropagation (by depth), effectiveness β€” per-session quality signals
GET /api/sessions?limit=NSessions with started_at, cost, metadata β€” to bucket sessions into time windows

Report Sections

1. Windowing

Split history into a baseline window (older) and a recent window (newer). Default: recent = last 30 days, baseline = the 30–90 day range before it. Use daily_events/daily_sessions for series metrics and GET /api/sessions?limit=N to assign sessions to each window by started_at.

2. Error Rate Regression

  • Recent error rate = APIError count / total events in the recent window (from event_types and daily_events, or per-session GET /api/events).
  • Compare to the baseline rate. Flag if recent is higher.
  • Report the absolute and relative change and which sessions contributed most APIError events.

3. Cache Hit Rate Regression

  • Cache hit rate = total_cache_read / (total_cache_read + total_input).
  • Compute for each window (per-window input/cache_read from session metadata or the pricing breakdown). Flag a falling hit rate β€” that means more uncached input tokens and higher cost.

4. Compaction Frequency Regression

  • Compaction frequency = Compaction events / session per window (from event_types / daily_events, confirmed via per-session GET /api/workflows/{id} compaction). Flag a rising rate β€” context is overflowing more often.

5. Cost-per-Session Regression

  • Cost-per-session = window total cost / window session count, using GET /api/pricing/cost overall and GET /api/pricing/cost/{id} for the sessions in each window. Flag a climbing value.

6. Verdict

Roll up which metrics regressed, rank by relative worsening, and name the most likely driver (e.g., cache hit rate fell β†’ cost per session climbed).

Output

  • A Markdown table: metric | baseline | recent | Ξ” | direction (β–² worse / β–Ό better) | verdict.
  • Tag each regressed metric πŸ”΄ (clear regression), 🟑 (mild/within noise), or 🟒 (improved).
  • Currency in USD to 4 decimals; rates as percentages to 2 decimals.
  • List the specific session IDs that contributed most to any regression.
  • End with the single highest-priority regression to address and a concrete next step.
  • Read-only: only report what the API returns; never fabricate baselines.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.