agentsclimarketplace

Data analysis

Skill kenantang/codex-and-claude-skills/collected-academic-research-skills/sources/lingzhi227__agent-research-skills/skills/data-analysis

Generated by Codex. Read with caution.

Install
npx -y skills add kenantang/codex-and-claude-skills --skill data-analysis

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Generate statistical analysis code with 4-round review. Select appropriate statistical tests, interpret results, and produce analysis reports with p-values, effect sizes, and confidence intervals. Use when analyzing experimental data for a paper.

SKILL.md

3.5 KB, 777 tokens by cl100k_base, as published. Nobody here has run it

Data Analysis

Generate rigorous statistical analysis code with multi-round review.

Input

  • $0 — Data source (CSV, JSON, pickle, or experiment logs)
  • $1 — Research goal or hypothesis to test

References

  • 4-round code review prompts: ~/.claude/skills/data-analysis/references/review-prompts.md

Scripts

Statistical summary and comparison

python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --compare method --metric accuracy --output summary.json
python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --describe

Detects data types, recommends tests, runs comparisons, outputs effect sizes and significance stars. Requires numpy, scipy.

Format p-values

python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --values "0.001 0.05 0.23" --format stars
python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --csv results.csv --column pvalue --format latex

Formats p-values with stars, LaTeX notation, or plain text. Stdlib-only.

Workflow

Step 1: Generate Analysis Code

Structure the code with these sections:

  1. # IMPORT — pandas, numpy, scipy, statsmodels, sklearn
  2. # LOAD DATA — Load from original data files
  3. # DATASET PREPARATIONS — Missing values, units, exclusion criteria
  4. # DESCRIPTIVE STATISTICS — Summary tables if needed
  5. # PREPROCESSING — Dummy variables, normalization
  6. # ANALYSIS — Statistical tests per hypothesis
  7. # SAVE ADDITIONAL RESULTS — Extra results to pickle

Step 2: 4-Round Code Review

  1. Round 1 — Code Flaws: Mathematical/statistical errors, wrong calculations, trivial tests
  2. Round 2 — Data Handling: Missing values, units, preprocessing, test choice
  3. Round 3 — Per-Table: Sensible values, measures of uncertainty, missing data
  4. Round 4 — Cross-Table: Completeness, consistency, missing variables

Step 3: Produce Results

  • Every nominal value must have uncertainty (CI, STD, or p-value)
  • Statistical tests must be appropriate for the data type
  • Results must match actual data — never hallucinate

Allowed Packages

pandas, numpy, scipy, statsmodels, sklearn, pickle

Statistical Test Selection

Data TypeTest
Two groups, normalIndependent t-test
Two groups, non-normalMann-Whitney U
Paired samplesPaired t-test / Wilcoxon
Multiple groupsANOVA / Kruskal-Wallis
CategoricalChi-square / Fisher's exact
CorrelationPearson / Spearman
RegressionOLS / Logistic / Mixed effects

Rules

  • Always report p-values for statistical tests
  • Account for relevant confounding variables
  • Use inherent package functionality (e.g., formula = "y ~ a * b" for interactions)
  • Do not manually implement available statistical functions
  • Access dataframes using string-based column names, not integer indices

Related Skills

Gives 0 of the 12 instructions most data analysis skills give in 777 tokens

Counted across 286 of the 286 authors here whose files we hold, read 2026-08-06

  • use excel formulas instead of hardcoded calculated valuesin 35 of 286, across 7 files
  • match existing template conventions when modifying filesin 35 of 286, across 7 files
  • document sources for all hardcoded valuesin 35 of 286, across 7 files
  • write minimal concise python codein 35 of 286, across 7 files
  • place all assumptions in separate assumption cellsin 32 of 286, across 5 files
  • apply industry-standard color coding to financial modelsin 31 of 286, across 5 files
  • format years as text stringsin 30 of 286, across 3 files
  • recalculate formulas using recalc.py after modificationsin 30 of 286, across 3 files
  • format negative numbers using parenthesesin 30 of 286, across 3 files
  • fix all identified formula errors before finishingin 27 of 286, across 1 file
  • use colorblind-safe palettesin 19 of 286, across 12 files
  • Name tests after the prevented bugin 13 of 286, across 8 files

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.