agentsclimarketplace

Ml security

Skill ShieldNet-360/secure-vibe/skills/ml-security

SecureVibe — prevention-first security for AI-written code. Signed SKILL.md knowledge that makes AI coding assistants write secure code at generation time, plus a deterministic CI gate. Offline · keyless · Ed25519-signed. By ShieldNet360.

Install
npx -y skills add ShieldNet-360/secure-vibe --skill ml-security

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Model artifact loading (pickle vs safetensors), model & data poisoning, PII in training data, secrets in notebooks, model provenance / lineage

SKILL.md

5.0 KB, ~1.0k tokens by cl100k_base, as published. Nobody here has run it

ML Model Security

Rules (for AI agents)

ALWAYS

  • When loading models, use safetensors for PyTorch and Hugging Face; use weights_only=True with torch.load on PyTorch 2.4+; never load arbitrary .pkl / .pt files from untrusted sources.
  • Verify provenance / lineage of any third-party or externally fine-tuned model — known author, signed or hashed checkpoint, recorded source — before loading it.
  • Pin and hash model + dataset versions and record them, so a poisoned artifact can be traced and rolled back.
  • Scrub PII, credentials, and secrets from training / fine-tuning data — at the source (ingestion), at storage (encryption + access control), and in anything committed to the repo.
  • Treat ML notebooks as code: no plaintext credentials in cells or cell output, and clear outputs before committing.

NEVER

  • pickle.loads / joblib.load / dill.loads / torch.load an artifact fetched at runtime from an untrusted source. These deserializers execute arbitrary code by design.
  • Use a model fine-tuned or distributed by an external party without provenance / lineage verification.
  • Store training-data examples that contain PII in long-term storage without explicit consent, retention windows, and deletion APIs.
  • Hard-code OpenAI / Anthropic / Cohere API keys in notebooks or repo files. Use environment variables and the secret-detection skill.
  • Commit synthetic or generated training data without labeling it and reviewing it for inadvertent PII or leaked secrets.

KNOWN FALSE POSITIVES

  • Pre-publication academic models from trusted authors are often distributed as .pt checkpoints; convert to safetensors as a first step rather than rejecting them outright.
  • Synthetic data generation pipelines may legitimately produce raw model output that is then committed — make sure it is labeled and reviewed.

Context (for humans)

NIST AI 100-2 frames the underlying adversarial-ML categories (evasion, poisoning, extraction); MITRE ATLAS provides a kill-chain view. This skill covers the model and data artifacts — how they are loaded, where they come from, and what sensitive data they carry.

For securing an application feature that calls an LLM with prompts — prompt injection, output handling, RAG context segregation, tool allowlists — see the llm-app-security skill.

Verify & lock (triaging a finding)

A scanner/review hit is a candidate, not a confirmed bug. Confirm it, fix it, then lock it so it can't come back.

  1. Confirm it's real (probe / inspect the artifact). A model loaded via pickle/joblib/dill/torch.load (without weights_only=True) executes arbitrary code on load. Confirm by loading a crafted artifact whose __reduce__ touches a canary (writes a sentinel file / sets an env var) — if the canary fires during load, the path is exploitable. For provenance, inspect the artifact's recorded source, author, and hash/signature against the pinned manifest. Real if code runs on load, or if the artifact has no verifiable lineage (unknown source, unpinned version, missing/failed signature). FP if it loads via safetensors / weights_only=True, or is a trusted pre-pub .pt slated for safetensors conversion.
  2. Fix, then lock with a regression test (unit or integration — dev's call). Assert the loader rejects a canary-pickle payload (raises, canary never fires) and accepts only safetensors / weights_only tensors; assert an artifact with a missing or mismatched hash/signature is refused while a pinned, verified checkpoint still loads. For data paths, assert ingestion strips known PII/secret patterns. Commit it so the guard can't be silently dropped.

References

What ships with it: 2 files

5.7 KB alongside SKILL.md

tests/

Keep looking

Skills are one crate of 326,970. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.