agentsclimarketplace

Performance monitor

Skill ulpi-io/plugin-marketplace/plugins/404kidwiz/skills/performance-monitor

A curated collection of 7,800+ agent skills for Claude Desktop, sourced from skills.sh

Install
npx -y skills add ulpi-io/plugin-marketplace --skill performance-monitor

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Expert in observing, benchmarking, and optimizing AI agents. Specializes in token usage tracking, latency analysis, and quality evaluation metrics. Use when optimizing agent costs, measuring performance, or implementing evals. Triggers include "agent performance", "token usage", "latency optimization", "eval", "agent metrics", "cost optimization", "agent benchmarking".

SKILL.md

3.4 KB, as published. Nobody here has run it

Performance Monitor

Purpose

Provides expertise in monitoring, benchmarking, and optimizing AI agent performance. Specializes in token usage tracking, latency analysis, cost optimization, and implementing quality evaluation metrics (evals) for AI systems.

When to Use

  • Tracking token usage and costs for AI agents
  • Measuring and optimizing agent latency
  • Implementing evaluation metrics (evals)
  • Benchmarking agent quality and accuracy
  • Optimizing agent cost efficiency
  • Building observability for AI pipelines
  • Analyzing agent conversation patterns
  • Setting up A/B testing for agents

Quick Start

Invoke this skill when:

  • Optimizing AI agent costs and token usage
  • Measuring agent latency and performance
  • Implementing evaluation frameworks
  • Building observability for AI systems
  • Benchmarking agent quality

Do NOT invoke when:

  • General application performance → use /performance-engineer
  • Infrastructure monitoring → use /sre-engineer
  • ML model training optimization → use /ml-engineer
  • Prompt design → use /prompt-engineer

Decision Framework

Optimization Goal?
├── Cost Reduction
│   ├── Token usage → Prompt optimization
│   └── API calls → Caching, batching
├── Latency
│   ├── Time to first token → Streaming
│   └── Total response time → Model selection
├── Quality
│   ├── Accuracy → Evals with ground truth
│   └── Consistency → Multiple run analysis
└── Reliability
    └── Error rates, retry patterns

Core Workflows

1. Token Usage Tracking

  1. Instrument API calls to capture usage
  2. Track input vs output tokens separately
  3. Aggregate by agent, task, user
  4. Calculate costs per operation
  5. Build dashboards for visibility
  6. Set alerts for anomalous usage

2. Eval Framework Setup

  1. Define evaluation criteria
  2. Create test dataset with expected outputs
  3. Implement scoring functions
  4. Run automated eval pipeline
  5. Track scores over time
  6. Use for regression testing

3. Latency Optimization

  1. Measure baseline latency
  2. Identify bottlenecks (model, network, parsing)
  3. Implement streaming where applicable
  4. Optimize prompt length
  5. Consider model size tradeoffs
  6. Add caching for repeated queries

Best Practices

  • Track tokens separately from API call counts
  • Implement evals before optimizing
  • Use percentiles (p50, p95, p99) not averages for latency
  • Log prompt and response for debugging
  • Set cost budgets and alerts
  • Version prompts and track performance per version

Anti-Patterns

Anti-PatternProblemCorrect Approach
No token trackingSurprise costsInstrument all calls
Optimizing without evalsQuality regressionMeasure before optimizing
Average-only latencyHides tail latencyUse percentiles
No prompt versioningCan't correlate changesVersion and track
Ignoring cachingRepeated costsCache stable responses

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.