agentsclimarketplace

Dataset quality audit

Skill zebbern/claude-code-guide/skills/dataset-quality-audit

Claude Code Guide - Setup, Commands, workflows, agents, skills & tips-n-tricks go from beginner to power user!

Install
npx -y skills add zebbern/claude-code-guide --skill dataset-quality-audit

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

3.9 KB, as published. Nobody here has run it

dataset-quality-audit

A data quality auditing tool that runs 12-dimension quality checks on tabular data, producing per-dimension scores (0–100), an overall grade, and actionable fix suggestions.

Capabilities

DimensionDescription
Missing ValuesCount and percentage of null/NaN values per column
Duplicate RowsNumber and percentage of fully duplicated rows
Type ConsistencyMixed types within a single column (e.g., numbers mixed with text)
Value Range / OutliersOutlier detection using the IQR method
Format ComplianceConsistency of date, email, phone number, and other formatted fields
Uniqueness ConstraintsWhether ID-type columns contain duplicates
Whitespace IssuesLeading/trailing spaces, empty strings, whitespace-only values
Constant ColumnsColumns with only a single unique value (zero information)
Distribution SkewnessWhether numeric columns have excessive skewness
Column NamingSpaces, special characters, or inconsistent casing in column names
Cardinality AnomaliesUnusually high or low number of unique values
Cross-Column ConsistencyLogical checks across columns (e.g., start date before end date)

Quick Start

# Basic quality check
python3 scripts/data_quality_checker.py data.csv

# Save report as JSON
python3 scripts/data_quality_checker.py data.csv --output report.json

# Specify ID columns (for uniqueness checks)
python3 scripts/data_quality_checker.py users.csv --id-columns "user_id,email"

# Specify date columns (for format checks)
python3 scripts/data_quality_checker.py orders.csv --date-columns "created_at,updated_at"

Detailed Usage

Basic Invocation

python3 scripts/data_quality_checker.py <data-file> [options]

Parameters

ParameterShortRequiredDefaultDescription
inputYesPath to input file (CSV/TSV/Excel/JSON)
--output-oNostdoutPath for the JSON report output
--id-columns-idNoAuto-detectComma-separated column names that should be unique
--date-columns-dcNoAuto-detectComma-separated column names containing dates
--sample-sNoAll rowsNumber of rows to sample (useful for large files)
--encoding-eNoutf-8File encoding

Output Format (JSON)

{
  "file": "data.csv",
  "rows": 10000,
  "columns": 15,
  "overall_score": 78.5,
  "grade": "B",
  "dimensions": {
    "missing_values": {
      "score": 85.0,
      "issues": [
        {"column": "age", "missing_count": 150, "missing_pct": 1.5, "suggestion": "Fill with median or mode"}
      ]
    },
    "duplicates": {
      "score": 95.0,
      "issues": [...]
    }
  },
  "top_suggestions": [
    "Column 'age' has 1.5% missing values — consider filling with the median",
    "Found 200 fully duplicated rows — consider deduplication"
  ]
}

Grading Scale

GradeScore RangeMeaning
A+95–100Excellent quality — ready for use as-is
A90–95Good quality — minor issues only
B80–90Moderate quality — recommended to fix before use
C60–80Poor quality — significant cleaning required
D40–60Very poor quality — many issues need attention
F0–40Essentially unusable — requires re-collection or major cleanup

Dependencies

  • Python 3.8+
  • pandas
  • numpy
pip install pandas numpy

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.