Evidence review
Skill Touchdown-Labs/inference-optimization-agent-pack/skills/evidence-review
Loadable systems-thinking skill pack for full-stack inference optimization.
npx -y skills add Touchdown-Labs/inference-optimization-agent-pack --skill evidence-reviewAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Review inference optimization claims against task success, p95/p99 latency, cost, cache behavior, retries, quality gates, and artifact-backed evidence.
SKILL.md
1.6 KB, 316 tokens by cl100k_base, as published. Nobody here has run it
Evidence Review
Use this when someone claims an optimization worked.
The goal is to decide whether the result is real, useful, and safe to keep.
Required Metrics
Ask for:
- cost per successful task, completed call, or accepted asset;
- success rate;
- quality score or human acceptance;
- p50/p95/p99 latency;
- retry rate;
- tool-call count;
- cache hit rate;
- model/provider cost;
- human rework;
- traffic shape;
- failure examples.
Evidence Artifacts
Look for:
- traces;
- logs;
- eval outputs;
- invoice or usage export;
- benchmark config;
- prompt/context versions;
- route decisions;
- cache keys and hit/miss logs;
- sandbox or tool logs;
- before/after samples.
Claim Labels
Label every conclusion:
- measured;
- source-backed;
- inferred;
- assumed;
- not_found.
Red Flags
- Token cost improved but success rate fell.
- Average latency improved but p99 got worse.
- A faster model caused more retries.
- A cache hit is counted without correctness or freshness checks.
- A self-hosted benchmark excludes operations cost.
- A media workflow reports render latency but not accepted asset rate.
- A voice workflow reports time-to-first-audio but not completed call outcome.
Output
Return:
- What is proven.
- What is not proven.
- What changed.
- Whether the change should ship, rollback, or run as an experiment.
- The next evidence packet needed.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most review quality skills give in 316 tokens
Counted across 1,048 of the 1,783 authors here whose files we hold, read 2026-08-07
- Ask questions one at a timein 81 of 1048, across 64 files
- Provide a recommended answer for each questionin 73 of 1048, across 50 files
- Explore the codebase instead of asking answerable questionsin 66 of 1048, across 42 files
- Resolve dependencies between decisions one-by-onein 42 of 1048, across 17 files
- Interview the user relentlessly about the planin 38 of 1048, across 13 files
- Order findings by severityin 31 of 1048
- Resolve each branch of the decision treein 27 of 1048, across 5 files
- Run a grilling sessionin 26 of 1048, across 5 files
- Update CONTEXT.md immediately when a term is resolvedin 26 of 1048, across 11 files
- Propose precise canonical terms for vague languagein 25 of 1048, across 7 files
- Create documentation files lazilyin 24 of 1048, across 5 files
- Assign severity to every findingin 24 of 1048
Said here and by no other author read
- ask for required metrics
- label every conclusion
- return what is proven
- return what changed
- state whether to ship rollback or experiment
- return the next evidence packet needed
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.