agentsclimarketplace

Auto skill build agent reliability loop

Skill Arnie016/codex-prompt-templates/skills/auto-skill-build-agent-reliability-loop

Color-coded Codex prompt templates and Agent Skills for plugin-orchestrated AI coding workflows, MCP safety, repo intelligence, and automation.

Install
npx -y skills add Arnie016/codex-prompt-templates --skill auto-skill-build-agent-reliability-loop

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Design production agent reliability loops across traces, evals, guardrails, gateways, model budgets, and shadow tests. Use when shipping AI/agent workflows that need simulate/evaluate/protect/monitor/optimize structure, gateway policy, eval suites, guardrail rollout, or trace-to-regression feedback. Skip when the task only needs a local Codex run ledger; use `$agent-observability` instead.

SKILL.md

2.9 KB, as published. Nobody here has run it

Agent Reliability Loop

Generated by: Codex Supercharge maintenance automation.

Goal: turn an agent workflow into a closed loop where production traces become eval cases, eval failures become guardrails or fixes, and gateway policy keeps cost, routing, and tool risk bounded.

Skip When

  • The user only wants local Codex run/cost tracking.
  • The work is a single benchmark or experiment with no production surface.
  • The task is only MCP conformance; use $auto-skill-build-mcp-conformance-harness.

Workflow

  1. Map the surface: list user journeys, models, tools, data classes, side effects, risk levels, and quality/cost/latency targets.
  2. Instrument first: require request IDs, trace/span IDs, session IDs, model/provider, token/cost, cache, fallback, guardrail, and tool metadata.
  3. Build eval bundles: create golden cases for task success, prompt conformance, tool correctness, unsafe requests, PII/secrets, latency, and cost. Keep judge prompts and heuristic checks versioned.
  4. Set gateway policy: define provider mapping, model fallbacks, cache rules, virtual keys, budgets, rate limits, privacy redaction, and audit logs.
  5. Stage guardrails: run new pre/post checks in log or monitor mode first, then enforce only after false positives and fail-open/fail-closed behavior are explicit.
  6. Shadow safely: mirror sampled traffic to candidate models or prompts only when shadow calls cannot trigger external side effects.
  7. Close the loop: cluster failed traces, match nearest successful traces, add representative failures to evals, patch prompts/tools/policies, then rerun the bundle before release.

Commands

rg -n "trace|span|cost|tokens|guardrail|fallback|budget|rate limit|request_id" .
rg -n "eval|rubric|judge|golden|dataset|experiment|shadow|canary" .
rg -n "tool|mcp|side effect|webhook|shell|filesystem|credential" .

Output

# Agent Reliability Plan
## Surface
## Instrumentation
## Eval Bundle
## Gateway Policy
## Guardrail Rollout
## Shadow Or Canary Plan
## Trace-To-Regression Loop
## Risks And Trust Notes
## Validation

Validation

  • Every production route has at least one trace and one regression case.
  • Every blocking guardrail has a false-positive review path.
  • Every budget/rate limit has an owner and an alert threshold.
  • Shadow/canary traffic cannot write to tools, accounts, payments, or user data.

Read references/future-agi-agent-reliability-loop.md for the source-backed pattern and risk notes.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.