agentsclimarketplace

Experiment log

Skill zakelfassi/skills-driven-development/examples/data-pipeline/skills/experiment-log

Agents that learn by doing — and remember how they did it. A methodology for AI agents to create, evolve, and share reusable skills.

Install
npx -y skills add zakelfassi/skills-driven-development --skill experiment-log

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 17 stars17 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Record a model training run — log parameters, metrics, and artifact paths, then append a structured row to the experiments ledger. Use when a training run completes, when asked to "log this experiment", or when comparing runs to decide which model to promote.

SKILL.md

3.6 KB, as published. Nobody here has run it

Experiment Log

Record a training run's parameters, metrics, and artifact paths in the experiments ledger.

Inputs

  • Experiment name (e.g., churn-v3-lgbm)
  • Run ID (auto-generated if omitted: {name}-{YYYYMMDD-HHmmss})
  • Parameters (key-value pairs, e.g., learning_rate=0.05 n_estimators=500)
  • Metrics (key-value pairs, e.g., auc=0.83 f1=0.71 precision=0.79 recall=0.64)
  • Artifact path (model checkpoint, serialized pipeline, etc.)
  • Notes (free-form, optional)

Steps

  1. Collect run metadata

    import datetime, hashlib, json, os
    
    run_id = "{name}-" + datetime.datetime.utcnow().strftime("%Y%m%d-%H%M%S")
    git_sha = os.popen("git rev-parse --short HEAD").read().strip()
    dataset_version = open("data/.version").read().strip()  # semver or hash
    
  2. Capture parameters and metrics If using a training script that outputs JSON:

    python train.py --config config/{name}.yaml --output-metrics /tmp/metrics.json
    

    Otherwise, capture values directly from the trainer object:

    params = model.get_params()
    metrics = {"auc": roc_auc_score(y_test, y_pred), "f1": f1_score(y_test, y_pred)}
    
  3. Save artifacts

    import joblib
    artifact_path = f"artifacts/{run_id}/model.pkl"
    os.makedirs(os.path.dirname(artifact_path), exist_ok=True)
    joblib.dump(model, artifact_path)
    
  4. Append to the experiments ledger

    scripts/log-experiment.sh \
      --name "{name}" \
      --run-id "{run_id}" \
      --params '{"learning_rate": 0.05}' \
      --metrics '{"auc": 0.83, "f1": 0.71}' \
      --artifact "{artifact_path}" \
      --notes "{notes}"
    

    The ledger is experiments/log.csv — a flat CSV with one row per run.

  5. Verify the entry

    tail -n 5 experiments/log.csv
    

    Confirm: run_id is unique, metrics are numeric, artifact path exists.

  6. Compare with previous runs (optional)

    python scripts/compare-runs.py --metric auc --top 5
    
  7. Promote if best (invoke pipeline-stage skill to wire the new model into serving) Only promote after peer review of the metrics and a sanity-check on the holdout set.

Conventions

  • Ledger location: experiments/log.csv — never delete rows; mark superseded runs with status=superseded
  • Artifact naming: artifacts/{run_id}/ — one directory per run
  • Required fields: run_id, experiment_name, timestamp, git_sha, auc (or primary metric), artifact_path
  • All metrics are floats; parameters are JSON strings
  • Every run references the dataset version it trained on (dataset_version column)

Edge Cases

  • Duplicate run_id: Append a counter suffix (-2, -3) rather than overwriting; log a warning.
  • Training failed mid-run: Log the entry with status=failed and the error message in notes; don't leave the ledger entry absent.
  • Very large artifact: Store only the path; add a artifact_size_mb column for reference.
  • Remote artifact store (S3/GCS): Use the full URI as artifact_path; the log script accepts both local paths and URIs.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.