Super ai ml ops
AI/ML operations: evaluation, monitoring, cost control, and reliability. Use for productionizing AI systems.From its SKILL.md
npx -y skills add arpitexplores/super-ai-ml-opsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
1.6 KB, 345 tokens by cl100k_base, as published. Nobody here has run it
Super AI/ML Ops
Overview
Make AI systems stable and measurable in production.
User Intent Examples
- "Need help with LLM Evaluation for my product/site."
- "Create a plan for LLM Ops."
- "Not sure where to start, need a quick assessment."
Workflow
- Define evaluation metrics, datasets, and acceptance thresholds.
- Set up observability for quality, latency, and errors.
- Implement caching and cost controls.
- Create monitoring and alerting for regressions.
- Establish release and rollback procedures for prompts/models.
- Document runbooks and ongoing QA cadence.
Minimal Intake Questions
- Primary goal or outcome
- Scope (pages, systems, teams, or timeframe)
- Constraints (tools, budget, timeline)
Output Format
- Eval plan and scoring rubric
- Monitoring and alerting checklist
- Cost and caching strategy
- Release and rollback plan
- Runbook and QA cadence
Routing Map (Modules)
- LLM Evaluation ->
references/modules/llm-evaluation.md - LLM Ops ->
references/modules/llm-ops.md
Bundled References
references/modules/scripts/assets/agents/
Compatibility Notes
- If any module references slash commands or tool-specific paths, translate them into plain-language steps.
- Keep outputs platform-agnostic unless the user specifies a specific tool, stack, or agent.
Guardrails
- Do not rely on single metrics; include qualitative checks.
- Track cost per request and cap budgets.
- Treat prompt/model updates as production changes.
What ships with it: 12 files
112.9 KB alongside SKILL.md
agents/
- openai.yaml183 B
examples/
- README.md528 B
references/
- modules/llm-evaluation.md73.0 KB
- modules/llm-ops.md26.7 KB
- CHANGELOG.md554 B
- .gitignore23 B
- INSTALL.md2.4 KB
- LICENSE1.0 KB
- product.json764 B
- PUBLISHING.md753 B
- README.md6.9 KB
- VERSION6 B
Gives 0 of the 12 instructions most ship operate skills give in 345 tokens
Counted across 779 of the 1,178 authors here whose files we hold, read 2026-08-07
- Document a rollback plan before deploymentin 41 of 779, across 22 files
- Update the changelogin 21 of 779, across 19 files
- Run the test suitein 20 of 779
- Create an annotated git tagin 20 of 779
- Clean up feature flags after full rolloutin 18 of 779, across 10 files
- Verify deployment health after launchin 18 of 779, across 10 files
- Test both feature flag statesin 17 of 779, across 9 files
- Verify the working tree is cleanin 17 of 779
- Make database migrations backward-compatiblein 16 of 779, across 8 files
- Set up error monitoring before launchin 15 of 779, across 7 files
- Monitor metrics at each rollout stagein 14 of 779, across 5 files
- Create a GitHub releasein 14 of 779
Said here and by no other author read
- Translate tool-specific paths into plain-language steps
- Keep outputs platform-agnostic unless specified otherwise
- Define evaluation metrics, datasets, and acceptance thresholds
- Set up observability for quality, latency, and errors
- Implement caching and cost controls
- Create monitoring and alerting for regressions
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.