Vps health check
Use when asked for a health check, status check, or "is everything up/healthy" on this VPS — checks system resources plus every running service (n8n, WAHA, code-server, Ollama, OpenClaw, Tailscale, Docker, Syncthing notes sync, VNC/XFCE4 desktop) and configured site liveness, and reports findings. Designed to run locally on the VPS (installed under its own `~/.claude/skills/`), not over SSH from another machine.From its SKILL.md
npx -y skills add tejask0/agent-skills --skill vps-health-checkAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
8.1 KB, ~2.3k tokens by cl100k_base, as published. Nobody here has run it
VPS Health Check
Read-only diagnostic sweep of the VPS and every service it runs. All commands are read-only — safe to run without confirmation. Report a concise table per section and call out anything that isn't ✅.
Configuration
This skill needs a few host-specific details that vary per VPS. Look them up in
~/.claude/skills-user-config.json under a "vps-health-check" object with keys:
notes_vault_dir— path (relative to$HOME) of the Syncthing-synced notes vault, e.g."Notes"syncthing_gui_url— full URL of the Syncthing GUI, if it's bound to a non-localhost address (e.g. a Tailscale IP) rather thanlocalhost:8384own_device_name— this device's name as it appears in Syncthing's device list, so it's excluded from the "remote devices" completion loopsite_url— a public site to liveness-check at the end (e.g. a personal site served from this VPS); optional, skip the check if unset
If the file/keys are missing, ask the user once, then save them under the
"vps-health-check" key (create the file/dirs if needed, preserving any other
keys already in it) so future runs don't need to ask.
The service list below (n8n, WAHA, code-server, Ollama, OpenClaw, VNC) reflects one possible VPS setup — adjust container/service names, ports, and systemd unit names to match whatever is actually running on this machine.
How to run
Run the four batches below in parallel (independent), then the app-layer batch.
Paths use $HOME so they work regardless of the login user.
1. System resources
uptime; free -h; df -h /; du -sh $HOME/* 2>/dev/null | sort -rh | head -10; nproc
# Top 5 processes by RAM (RSS) and by CPU — spot runaway/memory-hungry procs
echo "-- top 5 by MEM --"; ps -eo pid,comm,%mem,%cpu,rss --sort=-%mem | head -6
echo "-- top 5 by CPU --"; ps -eo pid,comm,%cpu,%mem,rss --sort=-%cpu | head -6
Flag: disk / > 85%, swap absent if memory pressure, load > cores,
any single process > ~25% MEM or pegging a core (sustained ~100% CPU).
2. Docker
docker ps -a --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"
docker network inspect shared_net --format '{{range .Containers}}{{.Name}} {{end}}'
docker inspect -f '{{.Name}} restarts={{.RestartCount}}' n8n waha code-server
docker images -f dangling=true -q | wc -l # >0 = prunable
Flag: any container not Up, restarts climbing, expected containers missing from shared_net.
3. systemd + networking
systemctl --failed --no-pager
systemctl is-active ollama
systemctl --user is-active openclaw-gateway
systemctl --user is-active vncserver
ss -tlnp | grep 5901 # VNC port — should show Xtigervnc listening
tailscale status | head
4. Syncthing notes sync
CFG=$HOME/.claude/skills-user-config.json
NOTES_VAULT_DIR=$(jq -r '.["vps-health-check"].notes_vault_dir // "Notes"' "$CFG")
ST=$(jq -r '.["vps-health-check"].syncthing_gui_url // "http://localhost:8384"' "$CFG")
OWN_DEVICE_NAME=$(jq -r '.["vps-health-check"].own_device_name // ""' "$CFG")
systemctl --user is-active syncthing
APIKEY=$(grep -oP '<apikey>\K[^<]+' $HOME/.local/state/syncthing/config.xml)
curl -s -H "X-API-Key: $APIKEY" "$ST/rest/db/status?folder=notes-vault" \
| jq -r '"folder: state=\(.state) files=\(.localFiles) errors=\(.errors)"'
CONNS=$(curl -s -H "X-API-Key: $APIKEY" "$ST/rest/system/connections")
while IFS=$'\t' read -r did name; do
connected=$(printf '%s' "$CONNS" | jq -r --arg d "$did" '.connections[$d].connected // false')
pct=$(curl -s -H "X-API-Key: $APIKEY" "$ST/rest/db/completion?folder=notes-vault&device=$did" | jq '.completion | floor')
printf "device: %-28s connected=%s completion=%s%%\n" "$name" "$connected" "$pct"
done < <(curl -s -H "X-API-Key: $APIKEY" "$ST/rest/config/devices" \
| jq -r --arg own "$OWN_DEVICE_NAME" '.[] | select(.name != $own) | [.deviceID, .name] | @tsv')
find "$HOME/$NOTES_VAULT_DIR" -name '*.sync-conflict-*' 2>/dev/null | wc -l # >0 = conflicts to resolve
# rclone is DISABLED (superseded by Syncthing). Confirm it stays off — the cron line must be commented:
crontab -l 2>/dev/null | grep -q '^[^#]*rclone' && echo "WARN rclone cron ACTIVE (should be disabled)" || echo "rclone cron disabled ✅"
Healthy: service active, folder state=idle errors=0, each remote device connected=True, conflict count 0, rclone cron disabled (if this VPS has migrated off rclone — skip that check otherwise).
5. Restic backup (object storage)
# Last backup log — should show "Backup complete" within last ~2 hours
tail -4 $HOME/.restic-backup.log 2>/dev/null || echo "WARN: no backup log found"
# Cron still active?
crontab -l 2>/dev/null | grep -q 'notes-backup' && echo "cron active ✅" || echo "WARN: cron job missing"
# Latest snapshot on the remote (confirms remote reachable + data exists)
source $HOME/.restic-env 2>/dev/null && $HOME/bin/restic snapshots --last 1 --no-lock 2>&1
Healthy: log ends with "Backup complete" ≤2 h ago, cron active, at least 1 snapshot listed.
6. App-layer endpoints
curl -s -o /dev/null -w "n8n %{http_code}\n" http://localhost:5678/healthz
curl -s -o /dev/null -w "waha %{http_code}\n" http://localhost:3000/ # 401 expected
curl -s -o /dev/null -w "code-server %{http_code}\n" http://127.0.0.1:9000/ # 302 expected
curl -s -w " ollama\n" http://localhost:11434/api/version
curl -s -o /dev/null -w "openclaw %{http_code}\n" http://127.0.0.1:18789/
SITE_URL=$(jq -r '.["vps-health-check"].site_url // empty' $HOME/.claude/skills-user-config.json)
[ -n "$SITE_URL" ] && curl -s -o /dev/null -w "site %{http_code}\n" --max-time 10 "$SITE_URL"
Expected results (healthy baseline)
| Service | Healthy signal |
|---|---|
| n8n | HTTP 200 on /healthz |
| WAHA | HTTP 401 on / (API-key gated = up); session check needs X-Api-Key header |
| code-server | HTTP 302 |
| Ollama | 200 + version JSON |
| OpenClaw | HTTP 200; Telegram in/out in journalctl --user -u openclaw-gateway |
| Tailscale | this device listed (peer offline = client device, not a VPS issue) |
| Syncthing | service active; folder state=idle errors=0; remote devices connected=True; 0 conflict files |
| Restic backup | log ends "Backup complete" ≤2 h ago; cron active; ≥1 snapshot on remote |
| VNC desktop | vncserver service active; port 5901 listening (Xtigervnc process) |
Configured site (site_url) | HTTP 200 |
Common gotchas
- WAHA 401 is healthy, not an error. To verify the WhatsApp session itself,
check
docker logs waha --tail 20for the lateststate—CONNECTED/WORKINGis good; stuck inPAIRING/SCAN_QR_CODEneeds a QR re-scan. - "99% reclaimable" in
docker system dfis shared base layers, not freeable — trustdocker images -f dangling=trueinstead. - Syncthing GUI may be bound to a non-localhost address (e.g. a Tailscale IP) —
curl/REST calls must use
syncthing_gui_urlfrom config or they connection-refuse. - A remote device showing
connected=Falseis usually fine if that device is just offline (laptop closed, phone asleep) — Syncthing is a hub, not a server outage. Only flag if a device that should be on is disconnected. - Completion stuck at 99.99% (1 item, few KB) is usually a per-device
.stignoremismatch on a workspace/state file (e.g..obsidian/workspace*.json), not a sync failure. - rclone, if present, is intentionally disabled on VPS setups that migrated to Syncthing — an active rclone cron is the anomaly to flag, not its absence.
Report format
One overall verdict line (HEALTHY / DEGRADED), then a section table, then a short "Watch items" list for anything non-✅. Don't propose fixes unless asked — this is diagnostic.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.