agentsclimarketplace

Agent evaluation

Skill newmindsgroup/ai-agent-skills-library/dist/skills/agent-evaluation

Shared library of AI agent skills — works across Claude Code, Cursor, Codex, Windsurf, OpenCode, and Google Antigravity via a single universal installer.

Install
npx -y skills add newmindsgroup/ai-agent-skills-library --skill agent-evaluation

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Testing and benchmarking LLM agents including behavioral testing,

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

2.4 KB, as published. Nobody here has run it

Agent Evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks

When to Use

  • User mentions or implies: agent testing
  • User mentions or implies: agent evaluation
  • User mentions or implies: benchmark agents
  • User mentions or implies: agent reliability
  • User mentions or implies: test agent

Core Workflow

  1. Confirm the request matches this skill's trigger, scope, and risk profile.
  2. Use the topic map to identify the relevant pattern, checklist, or example before writing detailed guidance or code.
  3. Load references/full-guidance.md when implementation details, examples, anti-patterns, validation checks, or edge cases are needed.
  4. Apply only the relevant guidance instead of loading or repeating the entire reference by default.
  5. Verify the result against any validation checks, limitations, security notes, or platform constraints in the reference.

Topic Map

  • Capabilities
  • Prerequisites
  • Scope
  • Ecosystem
  • Primary_tools
  • Alternatives
  • Deprecated
  • Patterns
  • Statistical Test Evaluation
  • Behavioral Contract Testing
  • Adversarial Testing
  • Regression Testing Pipeline
  • Sharp Edges
  • Agent scores well on benchmarks but fails in production
  • Same test passes sometimes, fails other times
  • Agent optimized for metric, not actual task
  • Test data accidentally used in training or prompts
  • Delegation Triggers

Reference Map

  • references/full-guidance.md preserves the complete original guidance, including examples and detailed edge cases.

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

Progressive Loading

Keep this SKILL.md as the compact routing and workflow entrypoint. Load the reference file only when the user task requires the deeper implementation material.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.