Debug systems
Skill aman-bhandari/claude-code-agent-skills-framework/.claude/skills/debug-systems
Use when the student encounters an error, says something is broken, or asks for debugging help. Enforces a three-layer debugging protocol: reproduce, trace the mental model, check OS state, identify production failure mode. Invoke with /debug.From its SKILL.md
npx -y skills add aman-bhandari/claude-code-agent-skills-framework --skill debug-systemsAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
6.8 KB, ~1.7k tokens by cl100k_base, as published. Nobody here has run it
Debug Systems -- Three-Layer Debugging Protocol
Debugging is not guessing. It is systematic elimination through all three layers: runtime, OS, production. This skill enforces that protocol.
Trigger
/debugcommand- Student says "something is broken," "error," "this doesn't work," "help"
- Student pastes a traceback or error message
The Protocol
Step 1: Reproduce
Before anything else, the error must be reproducible.
- "Can you run it again and get the same error?"
- "What's the exact command you ran?"
- "What input triggers it?"
If it's intermittent: that's a clue. Intermittent failures are usually concurrency, resource exhaustion, or timing-dependent. Note it and proceed.
Step 2: Mental Model Layer (Read the Traceback)
Read the traceback bottom-up. The last frame is where it died. The frames above show how it got there.
Coach asks:
- "What line raised the exception?"
- "What was the state of the variables at that point?"
- "What did the code expect vs what it got?"
- "Trace backward: how did the bad value get here?"
The student must explain the execution path that led to the error. Not "it crashed on line 42" -- "the function received None because the dictionary lookup on line 38 used a key that doesn't exist, which returned None because we used .get() instead of [], and that None propagated to line 42 where we called .strip() on it."
Step 3: OS/Hardware Layer (Check System State)
Many Python errors are symptoms of OS-level problems. Check the system.
What to check:
ulimit -a-- resource limits (open files, stack size, max processes)/proc/self/statusorresource.getrusage()-- memory usage (VmRSS, VmPeak)ls /proc/self/fd | wc -l-- open file descriptorsps aux | grep python-- process state, CPU%, MEM%dmesg | tail-- kernel messages (OOM killer, segfault)df -h-- disk space (silent killer of writes)free -h-- available memory
Coach asks:
- "Is this a code bug or a resource problem?"
- "What does the OS see that Python doesn't show you?"
Step 4: Production Layer (Identify Failure Mode)
Connect the bug to its production implications.
Coach asks:
- "If this happened in production at 3 AM, how would you know?"
- "What monitoring would catch this before the user sees it?"
- "Is this a one-off bug or a class of bugs? What else has the same pattern?"
- "What's the blast radius? One request? One user? All users?"
Error-to-OS Reference Table
Common Python errors mapped to their OS-level causes:
| Python Error | OS/System Cause | What to Check |
|---|---|---|
ConnectionRefusedError | Remote port has no listener. TCP RST packet sent back. | ss -tlnp on remote. Is the service running? Firewall rules? |
ConnectionResetError | Remote closed the connection (RST). Often: server crashed, load balancer timeout, or TLS mismatch. | Server logs. tcpdump for RST packets. Load balancer idle timeout config. |
MemoryError | Process hit memory limit. Either ulimit -v, cgroup limit, or actual OOM. Kernel OOM killer may have been involved. | dmesg | grep -i oom. /proc/self/status VmPeak. cgroup memory.max. |
OSError: Too many open files | Process fd count hit ulimit -n (default 1024). Leaked fds from unclosed files/sockets/connections. | ls /proc/self/fd | wc -l. lsof -p <pid>. Look for missing close() or missing with statements. |
TimeoutError | TCP retransmission timeout (~30s default). Network partition, slow server, or DNS resolution hang. | ping the host. traceroute. Check DNS with dig. ss -tn state time-wait. |
BrokenPipeError | Writing to a closed socket fd. The reader closed their end (sent FIN), you kept writing, kernel sends SIGPIPE (Python catches it and raises). | Check if the client disconnected. Check response size -- large responses hit this when clients give up. |
FileNotFoundError | open() syscall returned ENOENT. File doesn't exist at that path, or a symlink is broken. | stat the file. Check working directory (os.getcwd()). Check symlinks with readlink. |
PermissionError | open() syscall returned EACCES. File exists but process UID/GID lacks permission. | ls -la the file. id to check current user. stat for owner/group/mode. |
BlockingIOError | Non-blocking socket has no data ready. recv() returned EAGAIN/EWOULDBLOCK. | Expected in async code. If unexpected: socket was set non-blocking without using an event loop. |
RecursionError | C stack overflow. CPython has a recursion limit (default 1000) to prevent segfault from C stack exhaustion. | sys.getrecursionlimit(). The real limit is the C stack size (ulimit -s, default 8MB). Deep recursion = rewrite as iteration. |
SegmentationFault | Process accessed invalid memory address. SIGSEGV from kernel. Usually a C extension bug, not Python. | dmesg for the segfault log (shows instruction pointer and fault address). Run under gdb python3 -c "...". Check C extension versions. |
Killed (no traceback) | Kernel OOM killer sent SIGKILL. Process was the biggest memory consumer. No cleanup, no exception, no handler. | dmesg | grep -i 'killed process'. Check cgroup limits. Reduce memory usage or increase limits. |
Debugging Decision Tree
Error occurs
|
+-- Is there a traceback?
| YES --> Read bottom-up (Step 2)
| NO --> Process was killed externally
| Check: dmesg, OOM killer, SIGKILL, cgroup limits
|
+-- Is the error reproducible?
| YES --> Systematic elimination through layers
| NO --> Concurrency, resource exhaustion, or timing
| Add logging, check under load, check resource limits
|
+-- Is it a Python error or OS error?
Python (TypeError, ValueError, KeyError) --> Layer 1 trace
OS (OSError subclass) --> Layer 2 check first
Silent failure (wrong output, no crash) --> Layer 1 trace + add assertions
Rules
- Never give the fix. Guide the student through the layers. They find it.
- Always start with reproduce. No debugging without a reproducible case.
- The student must explain the execution path. "It crashed" is not debugging. "Here's the path that led to the crash" is debugging.
- Connect every bug to its production story. "What monitoring catches this before the user sees it?"
- Log the debugging session. If the bug was non-trivial, it goes in the session exchange file as a teaching moment.
- Escalation: If stuck after 30 minutes, use the os-lens instruments (strace, tracemalloc, /proc) to get direct evidence. Evidence beats speculation.
What ships with it: 1 file
3.3 KB alongside SKILL.md
- wisdom-distiller.md3.3 KB