agentsclimarketplace

Evals

Skill adriannoes/awesome-agentic-ai/cursor-claude-codex/skills/igoruehara-spec-driven/skills/evals

329 agent skills (Cursor, Claude Code & Codex), 5,380 OpenClaw skills, 201 ML notebooks, 7 textbooks, 52 research papers, 17 industry reports for PMs, Designers & Developers.

Install
npx -y skills add adriannoes/awesome-agentic-ai --skill evals

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Use to evaluate how faithfully the implementation matches the spec — runs the eval that checks whether each AC-N is covered by a task and referenced in a test, counts SPEC_DEVIATION and reports a score per feature. Trigger with /evals.

SKILL.md

1.6 KB, as published. Nobody here has run it

Translation note: Originally authored in Portuguese (pt-BR) by Igor Uehara (igoruehara/spec-driven, MIT). Translated to English by this hub to keep the repository language consistent. Original content unchanged in meaning; see the upstream repo for the pt-BR source.

Skill: Spec→code fidelity evals

Measures whether what was built reflects the spec — the quality metric of the agent's output. Two layers: deterministic (script) and judgment.

1. Deterministic layer

node scripts/eval-spec-fidelity.mjs .

Reports, per feature: total ACs, covered by a task, referenced in test/code, and open SPEC_DEVIATION. Fails (exit 1) if any AC has no task — broken traceability. Reference in test is a warning until the feature is implemented.

2. Judgment layer (the script does not catch this)

  • Does the test for each AC-N actually exercise the Given/When/Then — or does it just cite the ID in an empty test?
  • Does the implementation cover the edge cases and respect the spec's "Out of scope"?
  • Do the open SPEC_DEVIATION items have a resolution (fix the code or update the spec/ADR)?

Output

Score per feature + the gaps. It complements /validar (UAT for a single feature) with a portfolio fidelity view. The same eval runs in CI (esteira.yml) — here it is the judgment counterpart.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.