agentsclimarketplace

Super ai ml ops

Skill arpitexplores/super-ai-ml-ops

AI/ML operations: evaluation, monitoring, cost control, and reliability. Use for productionizing AI systems.From its SKILL.md

Install
npx -y skills add arpitexplores/super-ai-ml-ops

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

1.6 KB, 345 tokens by cl100k_base, as published. Nobody here has run it

Super AI/ML Ops

Overview

Make AI systems stable and measurable in production.

User Intent Examples

  • "Need help with LLM Evaluation for my product/site."
  • "Create a plan for LLM Ops."
  • "Not sure where to start, need a quick assessment."

Workflow

  1. Define evaluation metrics, datasets, and acceptance thresholds.
  2. Set up observability for quality, latency, and errors.
  3. Implement caching and cost controls.
  4. Create monitoring and alerting for regressions.
  5. Establish release and rollback procedures for prompts/models.
  6. Document runbooks and ongoing QA cadence.

Minimal Intake Questions

  • Primary goal or outcome
  • Scope (pages, systems, teams, or timeframe)
  • Constraints (tools, budget, timeline)

Output Format

  • Eval plan and scoring rubric
  • Monitoring and alerting checklist
  • Cost and caching strategy
  • Release and rollback plan
  • Runbook and QA cadence

Routing Map (Modules)

  • LLM Evaluation -> references/modules/llm-evaluation.md
  • LLM Ops -> references/modules/llm-ops.md

Bundled References

  • references/modules/
  • scripts/
  • assets/
  • agents/

Compatibility Notes

  • If any module references slash commands or tool-specific paths, translate them into plain-language steps.
  • Keep outputs platform-agnostic unless the user specifies a specific tool, stack, or agent.

Guardrails

  • Do not rely on single metrics; include qualitative checks.
  • Track cost per request and cap budgets.
  • Treat prompt/model updates as production changes.

What ships with it: 12 files

112.9 KB alongside SKILL.md

agents/

examples/

Gives 0 of the 12 instructions most ship operate skills give in 345 tokens

Counted across 779 of the 1,178 authors here whose files we hold, read 2026-08-07

  • Document a rollback plan before deploymentin 41 of 779, across 22 files
  • Update the changelogin 21 of 779, across 19 files
  • Run the test suitein 20 of 779
  • Create an annotated git tagin 20 of 779
  • Clean up feature flags after full rolloutin 18 of 779, across 10 files
  • Verify deployment health after launchin 18 of 779, across 10 files
  • Test both feature flag statesin 17 of 779, across 9 files
  • Verify the working tree is cleanin 17 of 779
  • Make database migrations backward-compatiblein 16 of 779, across 8 files
  • Set up error monitoring before launchin 15 of 779, across 7 files
  • Monitor metrics at each rollout stagein 14 of 779, across 5 files
  • Create a GitHub releasein 14 of 779

Said here and by no other author read

  • Translate tool-specific paths into plain-language steps
  • Keep outputs platform-agnostic unless specified otherwise
  • Define evaluation metrics, datasets, and acceptance thresholds
  • Set up observability for quality, latency, and errors
  • Implement caching and cost controls
  • Create monitoring and alerting for regressions

Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.

Keep looking

Skills are one crate of 326,851. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.