Data analyzer
Systematic exploratory data analysis. Activate when a dataset needs profiling — structure check, nulls, outliers, distributions, correlations — before deeper analysis begins.From its SKILL.md
npx -y skills add wachawo/claude-skills --skill data-analyzerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
2.3 KB, 481 tokens by cl100k_base, as published. Nobody here has run it
When to use
- You receive a new dataset and need to understand its shape and quality before analysis
- An analysis produces surprising numbers and you want to verify the underlying data first
- A stakeholder asks "is this data reliable?" or "what's in this table?"
- You're about to run a model or statistical test and need data-quality assurance
Process
- Load and overview — run
scripts/data_overview.pyto get row count, dtypes, memory usage, and a sample. Confirm grain (what one row represents). - Null profile — run
scripts/null_profiler.py; compare output against thresholds inreferences/quality_thresholds.mdand flag columns above limits. - Outlier detection — run
scripts/outlier_detector.py(IQR + z-score) on numeric columns; document flagged values and decide: real signal or data error? - Distribution summary — run
scripts/distribution_summary.pyfor descriptive stats and univariate histograms on each numeric column. - Correlation exploration — run
scripts/correlation_explorer.py; flag pairs with |r| > 0.8 as potential multicollinearity or redundancy. - EDA checklist sign-off — work through
references/eda_checklist.mdand confirm each item before declaring the dataset profiled. - Write findings — fill
assets/eda_report_template.mdwith full profiling output; distil top issues intoassets/findings_summary.md.
For pattern recipes (e.g. polars vs pandas equivalents, chunked reads for large files), see references/pandas_polars_recipes.md.
Inputs the skill needs
- Required: dataset path (CSV / Parquet / Excel) or a DataFrame already in scope
- Required: business context — what does one row represent?
- Optional: quality threshold overrides (defaults in
references/quality_thresholds.md) - Optional: columns to skip (PII, binary blobs, high-cardinality IDs)
Output
assets/eda_report_template.md(filled) — full profiling report with per-column statsassets/findings_summary.md(filled) — top 3–5 quality issues and recommended next steps- Console output / plots from scripts for interactive inspection
What ships with it: 10 files
25.1 KB alongside SKILL.md, 5 of them executable
assets/
- eda_report_template.md2.6 KB
- findings_summary.md1.4 KB
references/
- eda_checklist.md3.5 KB
- pandas_polars_recipes.md3.1 KB
- quality_thresholds.md2.3 KB
scripts/
- correlation_explorer.pyruns2.7 KB
- data_overview.pyruns2.2 KB
- distribution_summary.pyruns2.3 KB
- null_profiler.pyruns2.2 KB
- outlier_detector.pyruns2.8 KB