agentsclimarketplace

Validate the claim

Skill natexai2026/2030-skills/skills/validate-the-claim

Twelve portable reasoning skills for thinking clearly under uncertainty — a Claude Code plugin bundle (evidence grading, disaggregation, steelmanning, scenario branching, value frames, claim validation, hype checking, disparate-impact audit, anti-slop prose, and more).

Install
npx -y skills add natexai2026/2030-skills --skill validate-the-claim

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 29 days oldThe repository was created 29 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Separate what a vendor, lab, or benchmark reports about a system's performance from what has been independently validated in real conditions — and expose the gap before anyone relies on it. Use whenever evaluating a tool, model, product pitch, benchmark score, capability claim, or "our AI is X% accurate" number; during procurement, due diligence, or before deploying anything at scale. Triggers on "should we buy/deploy this," "is this benchmark real," "the vendor says," accuracy/uptime/win-rate claims, pilot proposals, and any impressive performance number whose source has an incentive to impress. If the claimant profits from you believing it, this skill applies.

SKILL.md

7.1 KB, ~1.5k tokens by cl100k_base, as published. Nobody here has run it

Validate the Claim

The most expensive mistakes come from trusting a number the seller produced. Vendor-reported performance and independently-validated performance are not the same measurement, and the gap between them can be enormous — a clinical model reported at 0.76–0.83 internally validated to 0.63 in the field, missing two-thirds of the cases it was sold to catch; 95% of enterprise AI pilots returning no measurable value against glossy case studies; a smart-building tool sold on "10–30% savings" where an industry survey found only 5% of pilots hit tangible results. This skill makes you ask the question the pitch is built to skip: validated how, by whom, in what conditions, against what?

The four gaps to interrogate

For any performance claim, work these four gaps in order. Each is a place the reported number and the real number quietly diverge.

1. Source gap — who produced this number?

  • Self-reported vs. independent. A vendor, lab, or team reporting on its own product has an incentive that shapes the metric. Treat self-reported numbers as tier-2-at-best and conflicted until independently checked (see grade-the-evidence).
  • Ask for the counterfactual. "Saved 853 employee-equivalents" against what baseline, measured how? Success stories rarely publish the comparison that would make the number meaningful.
  • Selection/survivorship. "80–90% Phase-I success for AI-discovered drugs" looks great until you notice the cohort is small, recent, and self-selected to advance only the winners. Impressive rates on a filtered sample are survivorship dressed as performance.

2. Benchmark-vs-deployment gap — was it measured where it'll be used?

  • Lab conditions ≠ field conditions. Biometric systems at "99.7% accuracy" in the lab degrade badly under uncontrolled lighting, motion, and adversarial input. Benchmarks "survive poorly on contact with real-world conditions."
  • Internal vs. external validation. Demand multi-site, external validation before scaling, not after. The internal-to-external drop is often where the whole value proposition dies.
  • Does the benchmark measure the real bottleneck? A model can beat experts on an exam whose skill isn't the actual constraint (the constraint is tacit skill, materials, integration, liability). High benchmark ≠ high real-world capability when the test omits the bottleneck.

3. Robustness gap — is it a fixed property or a function of effort?

  • Single-shot vs. sustained attack/use. "4.7% failure under one attempt" can become "63% at 100 attempts." Robustness is a statement about how hard someone tries, not a fixed number. Never cite a single-shot success rate as the verdict.
  • Overfitting tell. A large drop from the training/benchmark distribution to novel inputs (say, >10 points) signals reliance on brittle features, not robust capability. If a small perturbation breaks it, it wasn't solved.
  • Snapshot, not destiny. Defenders optimize against public benchmarks after publication. Today's score is a snapshot; don't project it forward as a durable ranking.

4. Failure-mode gap — what breaks, and does it self-correct?

  • Ask what the system does when it's wrong, and whether errors compound. Some systems degrade worse with more use (e.g. vulnerability density rising across successive AI code edits) — the opposite of the self-correction people assume.
  • Ask what's not measured. "Process metric moved" (a consult was ordered) is not "outcome improved" (the patient did better). Confirm the claim measures the thing you care about, not a proxy for it.

Procedure

  1. Pin the exact claim and its source. Who benefits if you believe it?
  2. Source gap: self-reported or independent? Counterfactual? Selection/survivorship?
  3. Deployment gap: lab or field? Internal or external validation? Does the benchmark hit the real bottleneck?
  4. Robustness gap: single-shot or sustained? Overfitting drop? Snapshot vs. durable?
  5. Failure-mode gap: what breaks, does it compound, is the metric an outcome or a proxy?
  6. State what independent evidence would justify reliance — and whether it exists yet. Give a deploy / pilot / don't-rely verdict.

Output template

CLAIM: <the performance claim> | SOURCE: <who> | THEY GAIN IF: <…>

SOURCE GAP:      self-reported | independent — counterfactual? selection?
DEPLOYMENT GAP:  lab | field — internal | external validation — bottleneck tested?
ROBUSTNESS GAP:  single-shot | sustained — overfitting drop? snapshot?
FAILURE-MODE GAP: what breaks / compounds? metric = outcome or proxy?

WHAT WOULD JUSTIFY RELIANCE: <the independent evidence needed — exists? no?>
VERDICT: deploy | pilot-with-instrumentation | don't-rely-yet — <why>

Example (compressed)

Claim: "Our AI triage model is 90% accurate — hospitals love it."

SOURCE GAP: vendor-reported. "90%" against no stated baseline; "hospitals love
  it" is testimonial, not measurement. Conflicted → tier 2 at best.
DEPLOYMENT GAP: the 90% is internal validation. No multi-site external number
  offered — which is exactly where comparable models have dropped to ~0.63 and
  started missing most true cases at the alert threshold.
ROBUSTNESS GAP: single operating point quoted; no sensitivity/specificity at the
  actual alert threshold, no false-alarm burden (which erodes clinician trust in
  all alerts if high).
FAILURE-MODE GAP: no data on what it misses or whom. "Accuracy" hides whether
  errors concentrate on a subgroup (→ hand to disparate-impact-audit).

WHAT WOULD JUSTIFY RELIANCE: independent, multi-site external validation
  reporting sensitivity/specificity at the deployed threshold, disaggregated by
  patient group. Does not exist here.
VERDICT: don't-rely-yet. Pilot only with independent instrumentation and a
  pre-registered external validation before any scale-up.

Gotchas

  • "Independent" must mean no incentive tie. A study the vendor commissioned is not independent, however credentialed the authors.
  • Absence of outcome data is itself a finding. "No study compares real outcomes with vs. without the tool" is a red flag, not a neutral gap — say so.
  • A real win is still a win. This skill isn't blanket cynicism; some claims survive all four gaps. Report those as validated and move on.
  • Pairs with grade-the-evidence (tiers the sources), hype-check (prevalence vs. impact), and disparate-impact-audit (whose errors the aggregate accuracy hides).

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.