agentsclimarketplace

Data scrub

Skill shakeebshaan/claude-code-quant-skills/skills/data-scrub

Claude Code skills, slash commands, and hooks tuned for quant research, backtesting, and crypto trading workflows.

Install
npx -y skills add shakeebshaan/claude-code-quant-skills --skill data-scrub

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Audit a market-data CSV for missing bars, timezone drift, weekend handling, stale ticks, duplicate timestamps, and survivorship. Use before running any backtest on new data.

SKILL.md

2.8 KB, 706 tokens by cl100k_base, as published. Nobody here has run it

Data Scrub Skill

Garbage in → backtest lies. Before ANY strategy runs on new data, audit it.

Ask for

  • File path
  • Expected timeframe (1m, 5m, 1h, 1d)
  • Expected timezone (UTC recommended)
  • Expected date range
  • Symbol(s) covered

Run these checks — each must PASS

1. Schema

  • Required columns: timestamp, open, high, low, close, volume
  • Dtypes correct? (timestamp = datetime64, OHLCV = float)

2. Timestamp integrity

  • Monotonic increasing, no duplicates
  • No gaps greater than 1 bar (for 24/7 crypto) or expected gaps (market hours for equities)
  • Timezone explicit and consistent — never naive
  • No daylight-saving jumps

3. Bar integrity

  • low <= open <= high for every bar
  • low <= close <= high for every bar
  • low <= high (sanity)
  • volume >= 0

4. Missing data

  • Count NaN in each column
  • Flag any bar where volume = 0 AND OHLC all equal → likely a stale / synthetic bar
  • Flag bars where high == low with non-zero volume → suspicious

5. Outliers

  • Return series: count bars with |return| > 20% (flag for review)
  • Volume spikes > 20x 30-bar rolling median → real or data error?

6. Symbol-level (if multi-symbol)

  • Symbol universe consistent across time, or point-in-time?
  • Delistings handled?
  • Ticker changes / splits adjusted?

7. Corporate actions (equities)

  • Dividend-adjusted?
  • Split-adjusted?
  • If not, warn user before using for backtests.

Code template

import pandas as pd

def scrub(df: pd.DataFrame, timeframe: str = "1h") -> dict:
    issues = {}
    expected_delta = pd.Timedelta(timeframe)

    issues["dupes"]       = df.index.duplicated().sum()
    issues["monotonic"]   = df.index.is_monotonic_increasing
    issues["gaps"]        = ((df.index.to_series().diff() > expected_delta).sum())
    issues["ohlc_break"]  = ((df["low"] > df["open"]) |
                             (df["low"] > df["close"]) |
                             (df["high"] < df["open"]) |
                             (df["high"] < df["close"]) |
                             (df["low"] > df["high"])).sum()
    issues["neg_volume"]  = (df["volume"] < 0).sum()
    issues["nan_rows"]    = df.isna().any(axis=1).sum()
    issues["stale_bars"]  = ((df["volume"] == 0) &
                             (df["open"] == df["close"])).sum()
    return issues

Output format

Single markdown table with Check / Result / Severity / Fix.

Severity:

  • BLOCK — do not backtest until fixed
  • WARN — acknowledge and document
  • OK — passed

End with a one-sentence verdict: ready for backtest / fixes required.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.