agentsclimarketplace

Verifying results before claiming

Skill K-Dense-AI/science-superpowers/skills/verifying-results-before-claiming

Composable computational-science methodology skills for AI research agents — pre-registration over TDD. A science-domain reimplementation of Superpowers.

Install
npx -y skills add K-Dense-AI/science-superpowers --skill verifying-results-before-claiming

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

What its author says it does

Copied from the file, not written here

Use when about to claim a result, effect, significance, or that an analysis reproduces, before reporting or writing it up - requires running the analysis fresh and reading the actual output first; evidence before claims always

SKILL.md

5.9 KB, ~1.4k tokens by cl100k_base, as published. Nobody here has run it

Verifying Results Before Claiming

Overview

Claiming a finding without fresh verification is dishonesty, not efficiency.

Core principle: Evidence before claims, always.

Violating the letter of this rule is violating the spirit of this rule.

The Iron Law

NO CLAIMS WITHOUT FRESH REPRODUCED EVIDENCE

If you haven't run the analysis in this state and read its actual output, you cannot claim its result. "It was significant earlier" is not evidence now.

The Gate Function

BEFORE claiming any result or expressing satisfaction:

1. IDENTIFY: What command/analysis proves this claim?
2. RUN: Execute it fresh and complete (from the immutable raw data, fixed seed)
3. READ: The actual output — the estimate, the interval, the p-value, the diagnostics
4. CHECK: Do the method's assumptions hold? Does it reproduce?
   For a CONFIRMATORY claim: does the pre-registration audit pass?
   (prereg.sh, ships with science-superpowers:preregistering-analysis:
    <skills root>/preregistering-analysis/prereg.sh audit)
5. VERIFY: Does the output actually support the claim?
   - If NO: state the real result with evidence
   - If YES: state the claim WITH the evidence (number + interval)
6. ONLY THEN: make the claim

Skip any step = asserting, not verifying

Common Failures

ClaimRequiresNot Sufficient
"The effect is significant"Fresh run; read estimate, CI, and p"It was significant before"
"There's no effect"Effect size + interval showing precisionA non-significant p (could be underpowered)
"The result reproduces"Re-run from raw data + fixed seed → same number"It ran fine earlier"
"Assumptions are met"The diagnostic output, read"It's probably fine"
"The model is good"Out-of-sample metricIn-sample fit / training accuracy
"Data cleaned correctly"Validation counts (rows in/out, ranges)"The script ran without error"
"Confirmatory finding"prereg.sh audit passes AND re-run as registeredThe prereg file was committed before the results (it may have been edited after)
"The subagent finished"Inspect the committed artifacts/diffThe subagent said "done"

Red Flags - STOP

  • Using "should", "probably", "seems to", "looks significant"
  • Expressing satisfaction before verification ("Great, it worked!", "Perfect!", "Confirmed!")
  • About to write up / report / commit a conclusion without a fresh run
  • Reporting a p-value you computed before the latest code change
  • Calling a result "reproducible" without having re-run it from raw
  • Trusting a subagent's success report
  • Calling a finding confirmatory without a passing prereg.sh audit — commit order alone is not a freeze
  • Describing a non-significant result as "no effect" without checking power/interval
  • ANY wording implying a finding without having just produced the evidence

Rationalization Prevention

ExcuseReality
"It was significant last run"Re-run it now. Code/data may have changed.
"I'm confident in the result"Confidence is not evidence.
"The script ran without errors"Running ≠ correct. Read the output.
"p < .05, so it's real"p is not the probability the effect is real. Report effect + interval.
"p > .05, so no effect"Could be underpowered. Absence of evidence ≠ evidence of absence.
"The subagent reported success"Verify the artifacts independently.
"The prereg was committed before the results, so it's frozen"Commit order misses post-freeze edits to the registration. Run prereg.sh audit; read the frozen-vs-current diff.
"It reproduces, I'm sure"Re-run from raw + seed and show the same number.
"Different words, so the rule doesn't apply"Spirit over letter.

Key Patterns

Significance / effect:

✅ [Re-run the model] [Read: beta=0.23, 95% CI [0.08, 0.38], p=.002] "Exposure raises the outcome; CI excludes zero."
❌ "The effect looked significant."

Reproducibility:

✅ Fresh env → re-run from data/raw with seed → same estimate to reported precision → "Reproduces."
❌ "It ran earlier, so it reproduces."

Null result:

✅ [Read: beta=0.01, 95% CI [-0.12, 0.14]] "No detectable effect; the interval rules out effects larger than ~0.14."
❌ "p=0.3, so there's no effect."

Model performance:

✅ [Held-out test set, used once] [AUC=0.78] "Out-of-sample AUC 0.78."
❌ "Training accuracy is 0.97, the model is great."

Confirmatory label:

✅ [Run prereg.sh audit] [Read: RESULT: PASS] → apply the frozen decision rule to the fresh output
❌ "The pre-registration predates the results commit, so the finding is confirmatory."

Subagent delegation:

✅ Subagent reports done → inspect committed code + output artifact → confirm → report actual state
❌ Trust the report

When To Apply

ALWAYS before:

  • Any statement of a result, effect, or significance
  • Any claim that an analysis reproduces or that assumptions hold
  • Any expression of satisfaction about a finding
  • Writing up, reporting, committing a conclusion, or requesting review
  • Trusting a subagent's output

Applies to: exact phrases, paraphrases, synonyms, and any implication of a verified finding.

The Bottom Line

Run it fresh. Read the output. Check it reproduces. THEN state the result — with the number and the interval.

This is non-negotiable.

Related Skills

  • science-superpowers:investigating-anomalous-results — if verification reveals the result doesn't hold
  • science-superpowers:requesting-red-team-review — independent scrutiny after you've verified

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 326,984. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.