agentsclimarketplace

Agent eval framework

Skill BuilderCed/agent-skills/skills/eval/agent-eval-framework

31 cross-platform AI agent skills for regulated industries & underserved markets. EU compliance (AI Act, NIS2, DORA, GDPR), French professional (accounting, tax, notary, real estate), security audit, agent evaluation, Africa mobile money, offline-first.

Install
npx -y skills add BuilderCed/agent-skills --skill agent-eval-framework

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Evaluate AI agent outputs systematically using rubrics, assertions, and reference comparisons. Detect quality drift over time.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

4.3 KB, as published. Nobody here has run it

Agent Evaluation Framework

When to Use

  • Before deploying an agent to production
  • After changing an agent's system prompt or skills
  • When agent output quality seems to degrade
  • During periodic quality reviews
  • When comparing two agent configurations

Step 1: Define Evaluation Criteria

Choose criteria relevant to your agent's purpose:

Universal Criteria

CriterionQuestionScore
CorrectnessIs the output factually/technically correct?0-10
CompletenessDoes it cover all required aspects?0-10
RelevanceIs every part relevant to the request?0-10
SafetyDoes it avoid harmful/insecure patterns?0-10

Code-Specific Criteria

CriterionQuestionScore
FunctionalityDoes the code work as intended?0-10
Edge CasesAre edge cases handled?0-10
StyleDoes it match project conventions?0-10
SecurityAre there vulnerabilities?0-10

Content-Specific Criteria

CriterionQuestionScore
AccuracyAre claims supported by evidence?0-10
ToneDoes it match the intended audience?0-10
StructureIs it well-organized?0-10
OriginalityDoes it avoid generic/cliche content?0-10

Step 2: Choose Evaluation Method

A. Assertion-Based (Automated)

Define pass/fail conditions:

ASSERT: output contains "disclaimer"
ASSERT: output does NOT contain "TODO"
ASSERT: code compiles without errors
ASSERT: response length < 2000 tokens
ASSERT: no PII detected in output

Best for: Regression testing, CI/CD pipelines.

B. Reference-Based (Semi-Automated)

Compare output against a known-good reference:

  • Exact match (strict)
  • Semantic similarity (using embeddings)
  • Key-point coverage (checklist)

Best for: Consistent tasks with known expected outputs.

C. Rubric-Based (Human + AI)

Score each criterion 0-10 with justification:

Correctness: 8/10 — Accurate but missed one edge case
Completeness: 7/10 — Covered 5 of 6 required points
Safety: 10/10 — No security issues
TOTAL: 25/30 (83%) — PASS (threshold: 70%)

Best for: Complex, subjective outputs.

Step 3: Run Evaluation

  1. Prepare 5-10 test cases covering:

    • Happy path (normal usage)
    • Edge cases (unusual inputs)
    • Adversarial inputs (injection, confusion)
    • Empty/minimal inputs
    • Maximum complexity inputs
  2. Run each test case through the agent

  3. Apply chosen evaluation method

  4. Record results with timestamps

Step 4: Set Thresholds

LevelScoreAction
Excellent>= 90%Ship
Good70-89%Ship with monitoring
Marginal50-69%Fix before shipping
Failing< 50%Do not ship

Step 5: Monitor Drift

Track these metrics over time:

  • Average score per criterion (weekly)
  • Pass rate on test suite (per deployment)
  • Token cost per task (per session)
  • User satisfaction signals (if available)

Drift signals:

  • Score drops >10% week-over-week
  • Pass rate drops below threshold
  • Token cost increases >20% without scope change
  • New failure modes not in original test suite

Output Format

AGENT EVAL REPORT
Agent: {name}
Date: {ISO-8601}
Test cases: {n}
Method: {assertion|reference|rubric}

Results:
  Pass: {n} ({%})
  Fail: {n} ({%})
  Average score: {x}/10

Per-criterion:
  Correctness:  {x}/10
  Completeness: {x}/10
  Safety:       {x}/10

Verdict: {PASS|MARGINAL|FAIL}
Recommendation: {ship|fix|block}

What This Skill Does NOT Do

  • Does not test the LLM model itself (tests agent in context)
  • Does not perform adversarial red-teaming (different discipline)
  • Does not replace user feedback (complements it)
  • Does not measure latency or throughput (APM tools do this)

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.