agentsclimarketplace

Evidence review

Skill Touchdown-Labs/inference-optimization-agent-pack/skills/evidence-review

Loadable systems-thinking skill pack for full-stack inference optimization.

Install
npx -y skills add Touchdown-Labs/inference-optimization-agent-pack --skill evidence-review

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Review inference optimization claims against task success, p95/p99 latency, cost, cache behavior, retries, quality gates, and artifact-backed evidence.

SKILL.md

1.6 KB, 316 tokens by cl100k_base, as published. Nobody here has run it

Evidence Review

Use this when someone claims an optimization worked.

The goal is to decide whether the result is real, useful, and safe to keep.

Required Metrics

Ask for:

  • cost per successful task, completed call, or accepted asset;
  • success rate;
  • quality score or human acceptance;
  • p50/p95/p99 latency;
  • retry rate;
  • tool-call count;
  • cache hit rate;
  • model/provider cost;
  • human rework;
  • traffic shape;
  • failure examples.

Evidence Artifacts

Look for:

  • traces;
  • logs;
  • eval outputs;
  • invoice or usage export;
  • benchmark config;
  • prompt/context versions;
  • route decisions;
  • cache keys and hit/miss logs;
  • sandbox or tool logs;
  • before/after samples.

Claim Labels

Label every conclusion:

  • measured;
  • source-backed;
  • inferred;
  • assumed;
  • not_found.

Red Flags

  • Token cost improved but success rate fell.
  • Average latency improved but p99 got worse.
  • A faster model caused more retries.
  • A cache hit is counted without correctness or freshness checks.
  • A self-hosted benchmark excludes operations cost.
  • A media workflow reports render latency but not accepted asset rate.
  • A voice workflow reports time-to-first-audio but not completed call outcome.

Output

Return:

  1. What is proven.
  2. What is not proven.
  3. What changed.
  4. Whether the change should ship, rollback, or run as an experiment.
  5. The next evidence packet needed.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Gives 0 of the 12 instructions most review quality skills give in 316 tokens

Counted across 1,048 of the 1,783 authors here whose files we hold, read 2026-08-07

  • Ask questions one at a timein 81 of 1048, across 64 files
  • Provide a recommended answer for each questionin 73 of 1048, across 50 files
  • Explore the codebase instead of asking answerable questionsin 66 of 1048, across 42 files
  • Resolve dependencies between decisions one-by-onein 42 of 1048, across 17 files
  • Interview the user relentlessly about the planin 38 of 1048, across 13 files
  • Order findings by severityin 31 of 1048
  • Resolve each branch of the decision treein 27 of 1048, across 5 files
  • Run a grilling sessionin 26 of 1048, across 5 files
  • Update CONTEXT.md immediately when a term is resolvedin 26 of 1048, across 11 files
  • Propose precise canonical terms for vague languagein 25 of 1048, across 7 files
  • Create documentation files lazilyin 24 of 1048, across 5 files
  • Assign severity to every findingin 24 of 1048

Said here and by no other author read

  • ask for required metrics
  • label every conclusion
  • return what is proven
  • return what changed
  • state whether to ship rollback or experiment
  • return the next evidence packet needed

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,984. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.