agentsclimarketplace

Dvc dataset versioning

Skill a5c-ai/babysitter/library/specializations/data-science-ml/skills/dvc-dataset-versioning

Dataset versioning skill using DVC for tracking data changes, managing data pipelines, and ensuring reproducibility.From its SKILL.md

Install
npx -y skills add a5c-ai/babysitter --skill dvc-dataset-versioning

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

SKILL.md

2.9 KB, 655 tokens by cl100k_base, as published. Nobody here has run it

dvc-dataset-versioning

Overview

Dataset versioning skill using DVC (Data Version Control) for tracking data changes, managing data pipelines, and ensuring reproducibility in ML workflows.

Capabilities

  • Dataset version tracking
  • Data pipeline definition and execution
  • Remote storage management (S3, GCS, Azure, etc.)
  • Reproducibility enforcement
  • Data lineage tracking
  • Experiment comparison with data versions
  • Cache management for large datasets

Target Processes

  • Data Collection and Validation Pipeline
  • ML Model Retraining Pipeline
  • Feature Store Implementation

Tools and Libraries

  • DVC
  • Git
  • Remote storage SDKs (boto3, google-cloud-storage, etc.)

Input Schema

{
  "type": "object",
  "required": ["action"],
  "properties": {
    "action": {
      "type": "string",
      "enum": ["init", "add", "push", "pull", "diff", "checkout", "run", "repro"],
      "description": "DVC action to perform"
    },
    "paths": {
      "type": "array",
      "items": { "type": "string" },
      "description": "File or directory paths to track"
    },
    "remote": {
      "type": "string",
      "description": "Remote storage name"
    },
    "revision": {
      "type": "string",
      "description": "Git revision for checkout/diff"
    },
    "pipeline": {
      "type": "object",
      "description": "Pipeline stage definition for run action"
    }
  }
}

Output Schema

{
  "type": "object",
  "required": ["status", "action"],
  "properties": {
    "status": {
      "type": "string",
      "enum": ["success", "error"]
    },
    "action": {
      "type": "string"
    },
    "trackedFiles": {
      "type": "array",
      "items": { "type": "string" }
    },
    "changes": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "path": { "type": "string" },
          "status": { "type": "string" },
          "hash": { "type": "string" }
        }
      }
    },
    "remote": {
      "type": "object",
      "properties": {
        "name": { "type": "string" },
        "url": { "type": "string" },
        "syncStatus": { "type": "string" }
      }
    }
  }
}

Usage Example

{
  kind: 'skill',
  title: 'Version training dataset',
  skill: {
    name: 'dvc-dataset-versioning',
    context: {
      action: 'add',
      paths: ['data/train.csv', 'data/test.csv'],
      remote: 's3-bucket'
    }
  }
}

What ships with it: 1 file

1.1 KB alongside SKILL.md

Keep looking

Skills are one crate of 326,059. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.