agentsclimarketplace

Data analysis

Skill JPeetz/agent-skills/data-analysis

The definitive collection of cross-platform Agent Skills. Compatible with Claude Code, Codex, Cursor, OpenClaw, Gemini CLI, Copilot, Hermes. Curated weekly. Higher quality than any alternative.

Install
npx -y skills add JPeetz/agent-skills --skill data-analysis

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Comprehensive data analysis agent skill for loading, cleaning, exploring, visualizing, and reporting on structured datasets. Supports CSV, JSON, Excel, and SQL data sources. Produces statistical summaries, correlation matrices, time series analysis, regression models, hypothesis tests, and publication-quality visualizations.

SKILL.md

12.2 KB, as published. Nobody here has run it

Data Analysis Agent Skill

Overview

A production-grade agent skill for end-to-end data analysis workflows. Use this skill when you need to load datasets, inspect structure, clean messy data, perform statistical analyses, create visualizations, or generate structured reports from tabular data.

When to Trigger

Activate this skill when the user asks to:

  • Analyze data: "Analyze this CSV", "What patterns do you see in sales.csv?", "Explore this dataset", "Summarize the data in users.json"
  • Create charts/graphs: "Make a bar chart of…", "Plot revenue over time", "Visualize the correlation matrix", "Show me a heatmap of…"
  • Find patterns: "Find trends in this data", "Is there a correlation between X and Y?", "Cluster customers from this data", "Detect anomalies in…"
  • Generate reports: "Create a report from survey_results.xlsx", "Summarize quarterly metrics", "Build a dashboard from sales data"
  • Clean data: "Clean this messy dataset", "Fix missing values in…", "Normalize these columns", "Deduplicate this CSV"
  • Statistical testing: "Run a t-test on group A vs B", "Check if this distribution is normal", "Perform regression analysis", "Calculate confidence intervals"

Near-Miss Negatives — Do NOT Trigger

  • Questions about database schema design without actual data (e.g., "What columns should my users table have?")
  • Questions about spreadsheet software UI (e.g., "How do I freeze a row in Google Sheets?")
  • General math / statistics theory questions without a dataset context (e.g., "Explain the central limit theorem")
  • Pure SQL query writing without a data-analysis intent (e.g., "Write a query to join three tables" — use a SQL skill instead)
  • Questions about ETL pipeline architecture or data engineering (e.g., "Design a data ingestion pipeline")

Step-by-Step Workflow

Follow these phases in order. Skip phases that don't apply (e.g., if data is already clean) but always state that you're skipping and why.

Phase 1: Load

Determine the data source and load it into a DataFrame.

CSV    → pd.read_csv(filepath, ...)
JSON   → pd.read_json(filepath, ...)
Excel  → pd.read_excel(filepath, sheet_name=...)
SQL    → pd.read_sql_query(query, connection)

Checklist:

  • Identify encoding (try UTF-8, Latin-1, detect automatically)
  • For CSV: inspect delimiter (comma, tab, semicolon), quote character
  • For Excel: list available sheets, load the right one
  • For JSON: handle nested structures with pd.json_normalize() if needed
  • For SQL: confirm read-only access, never run destructive queries
  • Load a sample first if the dataset is large (>100k rows)

Phase 2: Inspect

Understand what you're working with before touching anything.

df.shape          # rows × columns
df.info()         # dtypes, non-null counts, memory
df.head(10)       # first rows
df.tail(5)        # last rows
df.describe()     # numeric summary stats
df.describe(include='object')  # categorical summary
df.dtypes         # column types
df.columns.tolist()  # column names

Checklist:

  • Report shape: rows × columns
  • List all columns with their dtypes
  • Show summary statistics for numeric columns
  • Show value counts for low-cardinality categorical columns
  • Flag potential issues: wrong dtypes, placeholder values, suspicious zeros

Phase 3: Clean

Address data quality issues. Never modify the source file without prompting the user first. Work on a copy.

Common operations:

IssueApproach
Missing valuesdf.isnull().sum() → decide drop vs impute
Wrong dtypespd.to_numeric(), pd.to_datetime(), astype()
OutliersIQR method, Z-score, domain-specific thresholds
Duplicatesdf.duplicated().sum()df.drop_duplicates()
Inconsistent strings.str.strip(), .str.lower(), .str.replace()
Date parsingpd.to_datetime() with format or infer
NormalizationMin-max scaling, Z-score standardization
Categorical encodingOne-hot, label encoding for ML prep

Rules:

  1. Always work on df_clean = df.copy(), never mutate the original in-place without explicit user consent.
  2. Report every change: "Dropped 47 duplicate rows (2.3% of data)", "Imputed missing age values with median (142 cells)".
  3. Flag suspicious patterns even if you don't fix them: "Column 'salary' has 340 zero values — verify if these are legitimate."
  4. If a cleaning decision is irreversible, ask first.

Phase 4: Analyze

Apply appropriate analytical methods based on the question.

Exploratory Data Analysis (EDA):

  • Univariate: histograms, box plots, value counts per column
  • Bivariate: scatter plots, correlation coefficients, grouped means
  • Multivariate: pair plots, correlation matrix heatmap, PCA

Statistical Methods (see references/statistical-methods.md):

GoalMethod
Compare two groupsIndependent t-test, Mann-Whitney U
Compare 3+ groupsOne-way ANOVA, Kruskal-Wallis
Relationship between two continuous varsPearson/Spearman correlation
Predict continuous outcomeLinear regression, polynomial regression
Predict categorical outcomeLogistic regression
Check normalityShapiro-Wilk test, Q-Q plot
Detect time trendsMoving averages, decomposition, stationarity tests
Find clustersK-means, hierarchical clustering, DBSCAN
Reduce dimensionsPCA, t-SNE (visualization only)

Time Series specifics:

  • Set datetime index: df.set_index('date', inplace=True)
  • Resample: df.resample('M').mean()
  • Rolling windows: df['value'].rolling(7).mean()
  • Decomposition: trend, seasonal, residual

Phase 5: Visualize

Choose the right chart for the data and question. See references/visualization-patterns.md for the full guide.

Library selection:

  • Static, publication-quality: matplotlib + seaborn
  • Interactive, exploratory: plotly
  • Statistical plots: seaborn (box, violin, pair, joint, heatmap)

Quick reference:

Data TypeQuestionChart
Categorical × NumericCompare amountsBar chart, box plot
Numeric × NumericRelationshipScatter plot, line chart
Time × NumericTrend over timeLine chart, area chart
Categorical × CategoricalCross-tabulationHeatmap, stacked bar
DistributionShape of dataHistogram, KDE, violin
Part-to-wholeProportionsPie chart* (≤5 categories), treemap
Correlation matrixRelationshipsHeatmap
RankingsOrderHorizontal bar chart

*Pie charts: use only when ≤5 categories and values sum to a meaningful whole. Prefer bar charts otherwise.

Best practices:

  • Always label axes and add a title
  • Use accessible color palettes (avoid red-green for colorblind users)
  • Sort bar charts by value unless categories have a natural order
  • Add data source and date to chart footnotes
  • For interactive charts, include hover tooltips

Phase 6: Report

Synthesize findings into a structured report.

Report structure:

  1. Executive Summary — 2–3 sentences with the key finding
  2. Data Overview — source, shape, date range, columns
  3. Data Quality — issues found, actions taken
  4. Key Findings — bullet points with numbers, ranked by importance
  5. Visualizations — inline charts with captions
  6. Statistical Results — test statistics, p-values, effect sizes
  7. Limitations & Caveats — data gaps, assumptions, edge cases
  8. Recommendations — actionable next steps or further analysis

Output formats:

  • Quick answer: plain text summary in chat with key numbers
  • Detailed report: Markdown document with embedded charts
  • Dashboard: interactive HTML with plotly (offer if >5 charts)
  • Export: offer to save cleaned data and charts as files

Tool-Aware Implementation

Python Libraries

This skill assumes Python 3.9+ with the following libraries available. Check availability before use; install missing packages as needed.

# Core
import pandas as pd
import numpy as np

# Visualization
import matplotlib.pyplot as plt
import seaborn as sns

# Interactive
import plotly.express as px
import plotly.graph_objects as go

# Statistics
from scipy import stats
from scipy.stats import norm, ttest_ind, f_oneway, pearsonr, spearmanr

# Optional: machine learning
from sklearn.preprocessing import StandardScaler, LabelEncoder
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
from sklearn.linear_model import LinearRegression, LogisticRegression

Matplotlib setup for non-interactive environments:

import matplotlib
matplotlib.use('Agg')  # headless rendering

Plotly in notebooks vs scripts:

# In Jupyter/notebook environments:
import plotly.io as pio
pio.renderers.default = 'notebook'

# For saving to files:
fig.write_html('chart.html')
fig.write_image('chart.png')

Platform-Specific Notes

PlatformMatplotlib backendFile outputNotes
Claude CodeAggwrite to filesSave charts as PNG/HTML, display from disk
CodexAggwrite to filesSame approach
CursorAgg or interactivewrite to filesCan open HTML in preview
Gemini CLIAggwrite to filesSave charts, display paths
OpenClawAggwrite to filesUse canvas for HTML output
CopilotAggwrite to filesStandard file-based approach

Safety & Guardrails

  1. Never modify source data in-place — always create a copy or backup before transformations. Offer to save cleaned data as a new file.
  2. Flag data quality issues — don't silently fix problems. Report missing values, outliers, and type inconsistencies before and after cleaning.
  3. Statistical honesty — report p-values and effect sizes, not just "significant" or "not significant". Don't p-hack by running multiple tests without correction. Mention when sample sizes are too small for reliable inference.
  4. Privacy awareness — if a dataset appears to contain PII (emails, phone numbers, names), warn the user and suggest anonymization before analysis.
  5. Large file handling — for files >100MB, use chunked reading (chunksize parameter) or sample before full analysis. Warn about memory constraints.
  6. SQL safety — use read-only connections. Never run INSERT, UPDATE, DELETE, DROP, or ALTER. Use transactions or connection strings that enforce read-only mode.
  7. Deterministic results — set random seeds for reproducible analysis: np.random.seed(42).

Scripts

scripts/validate_dataset.py

PEP 723 compliant data quality validation script. Run with:

python scripts/validate_dataset.py path/to/dataset.csv
# or
python scripts/validate_dataset.py path/to/dataset.json
# or
python scripts/validate_dataset.py path/to/dataset.xlsx --sheet "Sheet1"

Produces a structured quality report covering missing values, outliers, type consistency, duplicates, and basic statistics. See script docstring for details.

References

  • data-cleaning-guide.md — Handling nulls, outliers, type coercion, deduplication, and normalization patterns.
  • visualization-patterns.md — Chart type selection guide, color best practices, accessibility considerations.
  • statistical-methods.md — Descriptive statistics, hypothesis testing, regression, correlation, confidence intervals, and when to use each method.

Gives 0 of the 12 instructions most data analysis skills give

Counted across 286 of the 286 authors here whose files we hold, read 2026-08-06

  • use excel formulas instead of hardcoded calculated valuesin 35 of 286, across 7 files
  • match existing template conventions when modifying filesin 35 of 286, across 7 files
  • document sources for all hardcoded valuesin 35 of 286, across 7 files
  • write minimal concise python codein 35 of 286, across 7 files
  • place all assumptions in separate assumption cellsin 32 of 286, across 5 files
  • apply industry-standard color coding to financial modelsin 31 of 286, across 5 files
  • format years as text stringsin 30 of 286, across 3 files
  • recalculate formulas using recalc.py after modificationsin 30 of 286, across 3 files
  • format negative numbers using parenthesesin 30 of 286, across 3 files
  • fix all identified formula errors before finishingin 27 of 286, across 1 file
  • use colorblind-safe palettesin 19 of 286, across 12 files
  • Name tests after the prevented bugin 13 of 286, across 8 files

Said here and by no other author read

  • follow the load, inspect, clean, analyze, visualize, report workflow
  • state the reason when skipping any workflow phase
  • report every data cleaning change made
  • work on a copy of the data
  • ask before making irreversible cleaning decisions
  • flag suspicious data patterns

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.