Troubleshooting
Public Agent Skills for Terminal and Kubernetes
npx -y skills add chaterm/terminal-skills --skill troubleshootingAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
What its author says it does
Copied from the file, not written here
性能问题排查
SKILL.md
4.1 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it
性能问题排查
概述
性能瓶颈定位、资源争用分析技能。
快速诊断
USE 方法
# Utilization, Saturation, Errors
# CPU
# 利用率
mpstat -P ALL 1
# 饱和度
vmstat 1 | awk '{print $1}' # 运行队列
# 错误
dmesg | grep -i "cpu"
# 内存
# 利用率
free -m
# 饱和度
vmstat 1 | awk '{print $7,$8}' # si/so
# 错误
dmesg | grep -i "oom"
# 磁盘
# 利用率
iostat -x 1 | awk '{print $NF}' # %util
# 饱和度
iostat -x 1 | awk '{print $10}' # avgqu-sz
# 错误
dmesg | grep -i "error"
# 网络
# 利用率
sar -n DEV 1
# 饱和度
netstat -s | grep -i "overflow"
# 错误
ip -s link
60 秒诊断
# 1. 系统负载
uptime
# 2. 内核消息
dmesg | tail
# 3. 系统统计
vmstat 1 5
# 4. CPU 统计
mpstat -P ALL 1 5
# 5. 进程 CPU
pidstat 1 5
# 6. 磁盘 IO
iostat -xz 1 5
# 7. 内存使用
free -m
# 8. 网络统计
sar -n DEV 1 5
# 9. TCP 统计
sar -n TCP,ETCP 1 5
# 10. 进程列表
top -bn1 | head -20
CPU 问题排查
高 CPU 使用
# 找出高 CPU 进程
top -c
ps aux --sort=-%cpu | head
# 查看进程线程
top -H -p PID
ps -T -p PID
# CPU 分析
perf top -p PID
perf record -g -p PID -- sleep 30
perf report
CPU 等待
# 查看 iowait
vmstat 1
iostat -x 1
# 找出 IO 进程
iotop
pidstat -d 1
上下文切换
# 系统级
vmstat 1 | awk '{print $12,$13}'
# 进程级
pidstat -w 1
pidstat -wt -p PID 1
内存问题排查
内存不足
# 查看内存使用
free -m
cat /proc/meminfo
# 查看进程内存
ps aux --sort=-%mem | head
smem -rs pss
# 查看缓存
slabtop
cat /proc/slabinfo
内存泄漏
# 监控进程内存
while true; do
ps -o pid,vsz,rss,comm -p PID
sleep 60
done
# 使用 valgrind
valgrind --leak-check=full ./program
OOM 分析
# 查看 OOM 日志
dmesg | grep -i "oom"
journalctl -k | grep -i "oom"
# 查看 OOM 分数
cat /proc/PID/oom_score
cat /proc/PID/oom_score_adj
磁盘 IO 问题
IO 瓶颈
# 查看 IO 统计
iostat -x 1
# 关键指标
# %util > 80%: 设备繁忙
# await > 10ms: 延迟高
# avgqu-sz > 1: 队列积压
# 找出 IO 进程
iotop -o
pidstat -d 1
磁盘空间
# 查看空间
df -h
df -i # inode
# 找大文件
du -sh /* | sort -rh | head
find / -type f -size +100M
# 找已删除但占用空间的文件
lsof | grep deleted
网络问题排查
连接问题
# 查看连接状态
ss -s
netstat -an | awk '/tcp/ {print $6}' | sort | uniq -c
# TIME_WAIT 过多
ss -tan state time-wait | wc -l
# 连接队列溢出
netstat -s | grep -i "overflow"
ss -ltn
带宽问题
# 查看流量
iftop
nethogs
sar -n DEV 1
# 查看连接带宽
ss -ti
延迟问题
# 网络延迟
ping target
mtr target
# TCP 延迟
ss -ti | grep rtt
常见场景
场景 1:系统变慢排查
#!/bin/bash
echo "=== 系统负载 ==="
uptime
echo "=== CPU 使用 ==="
mpstat 1 3
echo "=== 内存使用 ==="
free -m
echo "=== 磁盘 IO ==="
iostat -x 1 3
echo "=== 高 CPU 进程 ==="
ps aux --sort=-%cpu | head -5
echo "=== 高内存进程 ==="
ps aux --sort=-%mem | head -5
场景 2:应用响应慢
#!/bin/bash
PID=$1
echo "=== 进程状态 ==="
ps -p $PID -o pid,stat,pcpu,pmem,cmd
echo "=== 线程状态 ==="
ps -T -p $PID
echo "=== 打开文件 ==="
lsof -p $PID | wc -l
echo "=== 网络连接 ==="
ss -tnp | grep $PID | wc -l
echo "=== 系统调用 ==="
strace -c -p $PID -o /tmp/strace.out &
sleep 10
kill %1
cat /tmp/strace.out
排查清单
| 症状 | 检查项 |
|---|---|
| 系统慢 | load、CPU、内存、IO |
| 响应慢 | 网络、磁盘、锁 |
| 内存高 | 泄漏、缓存、交换 |
| IO 高 | 进程、队列、设备 |
常用工具
# 综合工具
htop, atop, glances, nmon
# CPU
top, mpstat, perf, pidstat
# 内存
free, vmstat, smem, pmap
# 磁盘
iostat, iotop, blktrace
# 网络
ss, netstat, iftop, tcpdump
Gives 0 of the 12 instructions most debug triage skills give in ~1.8k tokens
Counted across 839 of the 1,149 authors here whose files we hold, read 2026-08-06
- investigate root cause before proposing any fixin 102 of 839, across 65 files
- read error messages completelyin 90 of 839, across 48 files
- create a failing test case before fixingin 84 of 839, across 44 files
- reproduce the issue consistentlyin 82 of 839, across 40 files
- change one variable at a timein 82 of 839, across 42 files
- check recent changesin 74 of 839, across 35 files
- write the regression test before fixingin 74 of 839, across 36 files
- fix the root cause not the symptomin 60 of 839, across 43 files
- implement a single fix at a timein 59 of 839, across 20 files
- trace data flow backward to the sourcein 50 of 839, across 20 files
- remove all debug instrumentationin 49 of 839, across 13 files
- form a single hypothesisin 48 of 839, across 18 files
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.