agentsclimarketplace

Log driven diagnosis

Skill yeaight7/agent-powerups/plugins/debugging-diagnostics/skills/log-driven-diagnosis

Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more

Install
npx -y skills add yeaight7/agent-powerups --skill log-driven-diagnosis

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 6 stars6 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when debugging complex runtime failures, distributed systems, or issues where a local debugger cannot be attached.

SKILL.md

2.9 KB, 648 tokens by cl100k_base, as published. Nobody here has run it

Purpose

When you cannot step through code, logs are your only visibility. Be methodical about extracting signal from noise — never dump whole log files into context.

When to Use

  • Runtime failures where a debugger cannot be attached
  • Distributed systems where the failure spans services
  • Issues reproducible only from log evidence

Inputs

  • Log files or a log explorer covering the incident window
  • The incident timestamp and, if available, a request/trace ID

Workflow

  1. Time-bound the search. Never dump the whole log file — always grep for timestamps around the reported incident, or use tail:

    tail -n 200 app.log                          # most recent context
    grep -n "2026-06-06T14:0" app.log            # window around the incident
    awk '/14:02:00/,/14:05:00/' app.log          # bounded slice between timestamps
    
  2. Identify the request ID. If the system uses distributed tracing or request IDs, find the ID associated with the error, then search the log corpus for only that ID to trace the complete lifecycle of the failed request:

    grep -n "ERROR" app.log | head -5            # find the failing entry and its ID
    grep -n "<request-id>" app.log               # full lifecycle of that request
    # structured (JSON) logs:
    jq -c 'select(.request_id == "<request-id>")' app.log.json
    
  3. Look for preceding warnings. The ERROR log is usually just the final crash. The actual root cause is often a WARNING or unexpected INFO log that occurred milliseconds earlier (e.g., a connection retry failing, or an empty array being returned):

    grep -n -B 20 "<error-text>" app.log | grep -inE "warn|retry|timeout|empty"
    
  4. Add missing logs. If the logs do not provide enough visibility, your first action must be to add temporary logging to the application, reproduce the bug, and gather the new signals. Do not guess blindly if the logs are insufficient.

Output

  • The traced lifecycle of the failing request
  • The suspected root-cause log line(s), including any preceding warnings
  • Any temporary logging added (flagged for later removal)

Verification

  • Every search was time-bounded or ID-bounded — no full-file dumps in context
  • Full request lifecycle traced when an ID exists
  • Lines preceding the error inspected, not just the error line itself
  • Temporary logging added (and flagged for removal) where visibility was missing

Failure Modes

  • Error-line tunnel vision — reading only the final ERROR while the cause sits in an earlier WARNING.
  • Context flooding — dumping megabytes of log into the conversation instead of bounded slices.
  • Blind guessing — iterating on fixes when the honest move is adding logging and reproducing once more.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.