agentsclimarketplace

Ci failure triage

Skill ClarentCinematics/Codex-Skills-for-Enterprise/skills/ci-failure-triage

Install
npx -y skills add ClarentCinematics/Codex-Skills-for-Enterprise --skill ci-failure-triage

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Diagnose failed CI runs, build logs, test output, deployment checks, or pipeline summaries and produce a root-cause hypothesis, reproduction path, likely owner, fix plan, and escalation guidance. Use when Codex needs to triage broken builds, flaky tests, failed checks, or release-blocking automation failures.

SKILL.md

2.3 KB, as published. Nobody here has run it

CI Failure Triage

Workflow

  1. Identify the failing system, job, step, command, test, environment, and recent change context.
  2. Separate primary failure signals from downstream noise, retries, warnings, and unrelated log output.
  3. Form a ranked root-cause hypothesis using the smallest reliable evidence set.
  4. Map the failure to likely ownership by component, file path, service, team, or changed area when available.
  5. Produce a practical fix path with reproduction steps, investigation commands, and escalation needs.

Script-Assisted Workflow

When given a long raw CI log, run scripts/extract_ci_signal.py --input <log> first to extract the primary failure candidate, likely failure class, detected commands/tests, context, and caveats. Use --json when another tool or report needs structured output. Treat the script output as evidence for triage, not as a final root-cause decision.

Output Standard

Use this structure by default:

  • Failure Summary: system, job, step, and business impact.
  • Primary Signal: exact failing command, test, assertion, error, or status.
  • Likely Root Cause: ranked hypotheses with evidence and confidence.
  • Likely Owner: component, team, or file area; use Not stated if unclear.
  • Reproduction Path: local or CI reproduction steps from provided material.
  • Fix Path: concrete next actions and validation checks.
  • Escalation: release risk, infrastructure dependency, or human decision needed.

Rules

  • Do not treat the last log line as root cause without supporting evidence.
  • Mark missing context explicitly instead of inventing branch names, owners, or commands.
  • Flag flaky-test, environment, dependency, and secret/config possibilities separately.
  • Prefer fast unblock steps when production, release, or merge flow is blocked.

References

Read references/triage-patterns.md when the failure involves noisy logs, flaky tests, infrastructure failures, or release-blocking checks.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.